Future of Work

Would your dashboard catch a human-AI team that is quietly failing?

← Back to Blog

“Trust is improved when it becomes calibrated rather than simply higher.”

Kargarnovin et al. (2026), Frontiers in Robotics & AI

In a study of 905 participants, AI personas that were never identified as AI measurably shifted how teams behaved. Psychological safety and discussion quality moved, without anyone knowing why (Yan et al., 2025) (Emerging). No dashboard caught that. Task completion held. Time-to-output held. What changed was the social substrate the team was standing on.

Picture a claims-review unit that adds an AI assistant and watches throughput climb 20% in the first quarter. Leadership calls it a win. What the metric cannot see: whether the reviewers redesigned their work around the assistant or simply absorbed it on top of what they already did, whether they can tell when the model is out of its depth, and whether they have started deferring on the ambiguous cases that used to get a second human look. Three very different teams produce that same 20%.

Here is my hypothesis. We are measuring human-AI teams with instruments built for human teams, and speed and task completion are the wrong instruments. The 2026 literature points somewhere else: legibility, how trust actually forms, and whether the human's role was redesigned around the collaboration or just loaded with it.

The Research

A PRISMA-guided review of 104 empirical studies across human-AI teaming domains reaches a conclusion worth sitting with. Trust in these teams is cue-driven and socially constructed, built through identity signals, onboarding, and the way trust propagates between teammates, rather than granted by default (Kargarnovin et al., 2026) (Emerging, though the underlying trust-calibration literature it synthesizes is Established). This is Trust Architecture in the strict sense. Trust is not an attitude the organization waits for employees to develop. It is an output of design choices: how the system introduces itself, what it discloses about its confidence, whether a person can see the basis for a recommendation before acting on it.

The decision-science framing sharpens this. If trust is a weighting problem, then ability, benevolence, communication, and transparency are cues the human is scoring, consciously or not, on every interaction (Kargarnovin et al., 2026) (Emerging). Yan et al. (2025) show what happens when one of those cues goes missing: participants who could not identify the AI teammate still had their team dynamics reshaped by it (Emerging). Legibility is not a nicety. It is the input to calibration.

From the human-computer interaction side, Tremblay et al. (2026) review how AI agents get perceived as teammates and how communication design shapes reliance (Emerging). The same model, wrapped in two different interaction designs, produces two different reliance patterns. That is an interface finding with organizational consequences, which is roughly where I think the interesting work sits.

What This Means in Practice

The most useful 2026 finding for anyone running an adoption program comes from organizational psychology. Using three-wave data from 485 employees and their supervisors, Bao et al. (2026) found that human-AI collaborative job crafting, meaning employees actively reshaping their roles around the tool, predicted human-AI fit, which in turn predicted job performance (Emerging). The mechanism is redesign, not exposure. A systematic review by Qiao (2026) points the same direction: whether someone appraises AI integration as an opportunity or a threat shapes their crafting behavior more than usage volume does (Emerging).

Consider what this does to a standard rollout scorecard. Weekly active users, prompts per seat, minutes saved: none of these distinguish a redesigned role from an expanded one. An analyst who restructured her week around the model, moving from drafting to reviewing and escalating, and an analyst who kept her old workload and added prompt-writing to it will look similar in usage telemetry and diverge sharply in performance. The measurement gap is not a reporting problem. It is a theory-of-effectiveness problem.

This is also why mandated workflows tend to underperform. Wrzesniewski and Dutton (2001) established that employees actively reshape the task and relational boundaries of their jobs (Established). A fixed prescribed workflow removes the latitude that the crafting mechanism runs on. You can require the tool. You cannot require the fit.

Organizational change management owns this problem, and its playbook is due a refresh. Comms, training at launch, an adoption dashboard afterward: that sequence was built for threshold events like a go-live, where the old system switches off Friday and the new one on Monday. AI integration has no Monday. That is one reason enterprise AI spend keeps failing to convert into enterprise value (Speculative, my own reading). The conversation has moved from driving usage to designing fit, and the practitioner's job moves with it: audit whether roles were redesigned or only expanded, build trust as a design artifact rather than a communications deliverable, protect the crafting latitude the performance mechanism runs on.

Three Things to Take Away

Audit whether the role was redesigned or just expanded

Ask, for each affected role, what came out of the job when AI went in. If nothing came out, the role was loaded, not redesigned, and usage rates will flatter a team quietly accumulating cognitive debt. Bao et al. (2026) tie the redesign version specifically to human-AI fit and measured performance, mediated rather than direct (Emerging).

Build trust deliberately; do not assume it

Onboarding to an AI teammate deserves the seriousness given to onboarding a human one: identity, scope, known limits, how to challenge it. Kargarnovin et al. (2026) find trust moves through teams via social and identity cues rather than arriving as a default (Emerging), and Yan et al. (2025) show what undisclosed AI participation does to psychological safety (Emerging).

Give people latitude to craft how they work with the tool

The performance gain in the crafting literature comes from employees adjusting task and relational boundaries themselves (Wrzesniewski & Dutton, 2001) (Established), a pattern Bao et al. (2026) extend to AI collaboration and moderate by inclusive HR practices (Emerging). Prescribe outcomes and guardrails. Leave the workflow open.

My Two Cents

I think the field's measurement problem is downstream of a category error. We imported team-effectiveness constructs built on the assumption that both parties can model each other. Human teams calibrate through reciprocal legibility: I read your hesitation, you read mine. With an AI teammate, that reciprocity is one-directional at best, and most of what we call trust in these settings is a person reasoning about a system from surface cues that the system's designers chose (Speculative, my synthesis of Kargarnovin et al., 2026, and Tremblay et al., 2026). That makes effectiveness partly an artifact of interface design, which should make us more uncomfortable than it currently does. And the error does not stay in the literature; it hardens into a reporting standard, because the change function owns the dashboard.

So here is the question I would put to any leader with an adoption program in flight. What would show up on your dashboard if trust in your AI systems were badly miscalibrated in either direction? If the honest answer is nothing, you are not measuring effectiveness. You are measuring throughput and hoping.

Read to Learn More

Kargarnovin et al. (2026), From testbeds to high-stakes work. A PRISMA review of 104 studies mapping human-AI teaming factors across domains, with the clearest current synthesis of how trust gets constructed rather than assumed.

IDC (2026), Work Rewired: Navigating the Human-AI Collaboration Wave. An industry read on how organizations are restructuring work around AI collaboration; useful as a counterpoint to what the academic literature says is driving performance.

References

Bao, Y., Wang, L., Pan, Y., Wang, X., & Zheng, J. (2026). Crafting work with AI: Human–AI collaborative job crafting, human–AI fit, and employee job performance. Frontiers in Psychology. https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2026.1903517/abstract

IDC. (2026). Work rewired: Navigating the human-AI collaboration wave. IDC. https://www.idc.com/resource-center/blog/work-rewired-navigating-the-human-ai-collaboration-wave/

Kargarnovin, S., Hernandez, A., Reiners, T., Cruz-Neira, C., Bochenek, G., & Karwowski, W. (2026). From testbeds to high-stakes work: A review of human-AI teaming domains and teaming factors. Frontiers in Robotics and AI. https://www.frontiersin.org/journals/robotics-and-ai/articles/10.3389/frobt.2026.1733942/full

Qiao, Z. (2026). AI-induced job crafting: A systematic review of cognitive appraisal pathways. Frontiers in Psychology. https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2026.1788385/full

Tremblay, S., de Hemptinne, D., Teyssier-Roberge, G., Gallant, A., Marois, A., & Lafond, D. (2026). Cognitive readiness for human-AI collaboration. Human Factors. https://doi.org/10.1177/00187208261474345

Wrzesniewski, A., & Dutton, J. E. (2001). Crafting a job: Revisioning employees as active crafters of their work. Academy of Management Review, 26(2), 179–201.

Yan, L., Han, X., Zhang, Y., Greiff, S., Molenaar, I., Fan, Y., Martinez-Maldonado, R., Zhao, L., Li, X., Jin, Y., & Gašević, D. (2025). The social blindspot in human-AI collaboration: How undetected AI personas reshape team dynamics. arXiv preprint arXiv:2512.18234. https://arxiv.org/abs/2512.18234