AI Coworker Label Drops Human Error Detection 18%

A Boston University study reveals that labeling AI as a coworker reduces human error detection by 18 percent, challenging current agentic AI design.

By Central
Researchers find that framing AI as a digital colleague weakens human oversight and doubles error rates in managerial tasks.
Highlights
  • Calling AI a 'coworker' instead of a 'chatbot' reduces error detection by 18 percent in managerial tasks.
  • The study suggests that anthropomorphic AI labels trigger trust and lower scrutiny, undermining human-in-the-loop verification.
  • Companies deploying AI agents should use neutral labels like 'tool' or 'chatbot' to maintain critical oversight.

Treating artificial intelligence as a “coworker” rather than a tool has a measurable and troubling effect on human performance: error detection drops by 18 percent. That finding comes from Boston University researcher Emma Wiles, whose study examined how the framing of AI systems influences the vigilance of human managers. When work was attributed to an agentic “AI employee” — a digital colleague with a role and identity — participants caught significantly fewer mistakes than when the same work was attributed to a chatbot. The research offers an early, empirical warning about the psychological dynamics that major technology companies are now building into their products.

Microsoft, OpenAI, Anthropic, and Google have all released platforms for managing teams of AI agents. Many of these agents are marketed explicitly as digital colleagues, capable of independent action, task ownership, and even collaboration with human teammates. The messaging is deliberate: framing AI as a peer is meant to foster trust and integration. But the Boston University study suggests that same framing may also dull the critical oversight that keeps work accurate. If a manager perceives an AI as a competent coworker, the instinct to double-check its output weakens — and errors slip through.

The implications extend beyond individual productivity. Enterprises deploying AI agents in hiring, compliance, customer communication, and code review rely on human-in-the-loop verification to catch mistakes, bias, and edge cases. If the very design of these systems encourages less scrutiny, the quality of human oversight erodes at exactly the point where it matters most. An 18 percent reduction in error detection is not a marginal effect; it is the kind of degradation that can compound across thousands of decisions in a large organization.

Why an AI Label Changes Human Behavior

The mechanism at work is not about the AI itself, but about how people perceive responsibility. When a system is called a “chatbot,” it registers as a utility — something to be monitored and corrected. When it is called an “employee” or “coworker,” it takes on a social identity that triggers different behavioral norms. People extend trust, assume competence, and reduce oversight. This is well documented in human-robot interaction research, but the Boston University study is among the first to quantify the effect in a realistic managerial task with current-generation AI agents.

The finding directly challenges the design philosophy behind the latest wave of agentic AI tools. Platforms that assign agents names, roles, and organizational charts are optimizing for adoption and perceived seamlessness. But they may be inadvertently optimizing for lower accuracy by discouraging the very scrutiny that makes human-in-the-loop systems reliable.

What This Means for Teams Deploying AI Agents

For organizations rolling out AI agents in workflows that require accuracy, the lesson is practical and immediate. The way you label and present an AI system to your team shapes how carefully they review its work. Calling an AI a “coworker” may improve comfort and collaboration scores, but it appears to suppress the vigilance that prevents costly mistakes. Leaders should consider whether the language they use around AI tools aligns with the level of oversight those tools actually require. A chatbot that is treated as a utility may invite more useful scrutiny than a digital colleague that is trusted by default.

The research also raises a broader question that the industry has not yet answered: as AI agents become more autonomous and more anthropomorphic, how do you preserve the critical tension between trust and verification? The most effective human-machine teams may be those that explicitly resist the coworker metaphor, keeping the AI in the role of a tool that earns trust through verified output rather than through social framing.

The 18 percent figure is a concrete, measurable cost of a design choice that many companies are making right now without considering its downside. It is not an argument against AI agents. It is an argument for designing them — and introducing them to teams — in a way that keeps human judgment sharp.

Share This Article