AI Assistance Benefits Differ Based on User Expertise

New study from MIT, Stanford, and Columbia finds that AI explanations can benefit non-experts but hinder clinicians in medical diagnosis.

By Central
The study highlights how AI assistance effects vary by user expertise, with clinicians performing best without explanations.
Highlights
  • Non-experts became more accurate when assisted by AI but deferred to the model even when it was wrong.
  • Clinicians performed best when given only the AI prediction without any accompanying explanation.
  • The findings challenge the assumption that more explanation always improves AI-assisted decision-making.

Artificial intelligence systems designed to assist with disease diagnosis are not one-size-fits-all tools. A new study from researchers at MIT, Stanford University, and Columbia University reveals that the benefits of AI assistance—and the risks of AI-generated explanations—depend heavily on the user’s level of medical expertise. The findings, published today in Nature Medicine, challenge a prevailing assumption in healthcare AI: that more explanation is always better, and that a well-performing model will naturally improve outcomes for all users.

In controlled experiments, non-experts and primary care providers were asked to diagnose skin diseases from medical images, with and without the help of different explainable AI systems. The results were striking. Non-experts became more accurate when assisted by AI, but primarily because they deferred to the model—trusting its predictions even when the model was wrong. Clinicians, by contrast, were not misled by incorrect AI assistance and performed best when given only the model’s prediction, with no accompanying explanation at all.

“Good AI systems can improve performance in some health settings, but this has to be balanced carefully with algorithmic deference that can lead to more error,” said Marzyeh Ghassemi, an associate professor in MIT’s Department of Electrical Engineering and Computer Science, a member of the Institute for Medical Engineering and Science, and a principal investigator at the Laboratory for Information and Decision Systems and the Abdul Latif Jameel Clinic for Machine Learning in Health. “We know that both AI and explainability methods can engage automation bias in humans, and this anchoring effect is something that must be accounted for when we design AI systems.”

AI Explainability Methods Produce Divergent Outcomes for Different User Groups

The study tested four distinct explainable AI approaches alongside a baseline condition where users received only the AI prediction and confidence level with no explanation. The four methods included a similar-image approach that reinforced the prediction by showing comparable cases, a heat-map technique that highlighted the most diagnostically relevant regions of an image, and a large language model (LLM) that explained the model’s reasoning in plain language.

Non-experts were asked to determine whether an image of a skin mole was cancerous. Clinicians faced a more complex task: providing a differential diagnosis of dermatological disease. The researchers found that all explainable AI approaches improved the accuracy of non-experts, but the gains were driven almost entirely by better detection of non-cancerous moles. When the AI model was correct, non-experts benefited. When it was wrong, their performance suffered disproportionately.

“But the reason non-expert users are better is because they are more reliant on the models,” Ghassemi said. “When the model is wrong, it hurts performance more than it helps performance when the model is right. We were just able to train very good AI models for this setting.”

The deference effect was largest with LLM-based explanations. Non-experts found these explanations more convincing when they were vague or generic, and they reported higher confidence in their answers even when those answers were wrong. The plausible, confident-sounding rationale of an LLM pulled non-experts toward incorrect conclusions, a phenomenon the researchers attribute to automation bias—the tendency to over-trust automated systems.

Clinicians, on the other hand, showed remarkable resilience to incorrect AI explanations. Of all the explainability methods tested, LLMs boosted clinician accuracy the least. The reason, the researchers found, lies in how each group uses the explanation. Clinicians already have a diagnostic hypothesis in mind. They check the AI’s output against their own training and expertise, so a bad explanation gets caught. Non-experts, lacking that independent foundation, use the explanation to form their initial opinion, making them vulnerable to any plausible-sounding rationale.

“It really comes down to how each group uses the explanation,” said lead author Orson Xu, an assistant professor in the Department of Biomedical Informatics at Columbia University. “A clinician already has a diagnosis in mind and checks the AI against their own training, so a bad explanation gets caught. Meanwhile, a non-expert can use that exact same explanation to form an opinion in the first place, so a plausible, confident-sounding rationale can pull them toward the wrong answer. The same tool ends up being an asset for one user and a liability for another.”

What Is the Deference Effect in AI-Assisted Diagnosis?

The deference effect describes the tendency of users to over-rely on AI recommendations, accepting them uncritically even when the model is incorrect. In this study, the researchers found that non-expert users who were most deferential to AI assistance were the worst performers on the diagnostic task without AI help. When an AI model made an error, these users lacked the domain knowledge to detect the mistake and instead followed the flawed recommendation with increased confidence, particularly when the AI provided a detailed but incorrect LLM-generated explanation.

The time at which users received the AI explanation also influenced their behavior. When the explanation was presented first, before the user had a chance to perform the diagnosis independently, deference increased significantly. This anchoring effect meant that users anchored their thinking to the AI’s explanation and struggled to deviate from it, even when the image contained features that contradicted the model’s prediction.

Fairness-Constrained Models Reduce Diagnostic Disparities Based on Skin Tone

One of the more encouraging findings from the study involved the use of a fairness-constrained model designed to combat bias against darker skin tones. When researchers deployed this model, it significantly improved overall accuracy and reduced diagnostic disparities based on skin tone among non-expert users. This result suggests that technical interventions at the model level can address known biases in dermatology AI, where training datasets have historically underrepresented darker skin tones.

The fairness-constrained model did not eliminate the deference effect, but it did reduce the harm caused by overreliance on flawed AI. When the model itself is less biased, the consequences of uncritical acceptance are less severe. This finding underscores a broader point: improving the underlying model is a necessary step, but it is not sufficient on its own. The way AI recommendations are presented to users, and the timing of that presentation, remain critical determinants of real-world outcomes.

Roxana Daneshjou, a co-author of the study and an assistant professor of biomedical data science and dermatology at Stanford University, highlighted the particular vulnerability of patients who increasingly turn to AI for self-diagnosis. “These findings are important as patients increasingly turn to AI to help with their health care,” Daneshjou said. “Our findings show that those with the least medical knowledge are most likely to be led astray when explainable AI models give an erroneous output.”

The Structural Advantage of Domain Expertise in AI-Assisted Decision-Making

The study reveals a fundamental asymmetry in how different user groups interact with AI explanations. Non-experts approach the diagnostic task with no pre-existing hypothesis. The AI explanation becomes the primary source of information, shaping their understanding from the ground up. Clinicians, by contrast, bring years of training and pattern recognition to the task. They use the AI as a second opinion, comparing its output to their own assessment rather than adopting it wholesale.

This structural difference has profound implications for how AI systems should be designed. For non-experts, the goal should be to encourage critical thinking and independent evaluation rather than passive acceptance. Simply providing more detailed explanations, as LLMs do, can backfire by making the AI’s output seem more authoritative than it actually is. For clinicians, the goal should be to provide the model’s best prediction without unnecessary cognitive load—explanations that experts do not need and that may slow them down or introduce noise.

“We really want AI to improve creativity and either upskill or fill in gaps where users are missing subtle presentations,” Ghassemi said. “Otherwise, we risk engaging automation bias and then, when the model is wrong, users can’t recover.”

How Should AI Systems Be Redesigned to Reduce Automation Bias?

The researchers propose several design changes that could mitigate the deference effect and improve outcomes for all users. Rather than presenting an AI explanation first, systems could require users to provide their own diagnostic hypothesis before seeing the AI’s recommendation. This forced reasoning step would give users a chance to form an independent judgment, making them less susceptible to anchoring on the model’s output.

Another approach involves reframing AI explanations as prompts for further consideration rather than definitive justifications. Instead of an LLM saying “this image shows melanoma because of irregular borders and asymmetry,” the system could say “consider melanoma among other possibilities—here are the features that support that diagnosis.” This subtle linguistic shift could encourage users to treat the AI as a collaborator rather than an oracle.

The study also found that AI systems outperformed humans when the presentation of disease was subtle, while humans performed much better when images contained atypical symptoms or unrelated features. This asymmetry suggests an opportunity for hybrid systems that dynamically adjust their level of explanation and deference based on the difficulty of the case and the known capabilities of the user.

For non-experts, a system that flags uncertainty explicitly—”I am not confident about this prediction”—may be more effective than one that always provides a confident-sounding explanation. The researchers note that LLM-based explanations, in particular, tend to be overconfident and overly detailed, creating a false sense of reliability that non-experts cannot easily evaluate.

Implications for Healthcare AI Deployment and Regulation

The study arrives at a time when FDA-approved AI interfaces are increasingly being used to help clinicians identify skin conditions in medical images. Several of these tools already include explainability features, and the market for AI-powered diagnostic support is growing rapidly. At the same time, non-experts are using AIa-powered search engines and consumer health apps that predict skin diseases based on user prompts, often with LLM-based explanations.

The research suggests that current regulatory frameworks may need to account not only for the accuracy of AI models but for the interaction between model output and user expertise. A model that performs well in isolation may perform poorly in practice if its explanations lead non-experts astray. Conversely, a model with modest standalone accuracy could be highly effective if its output is presented in a way that encourages critical evaluation.

“It’s getting obvious that we cannot just assume a good AI will solve all problems,” Xu said. “We need to pay careful attention to the users who will be using the AI system, because the same explanation can help an expert and mislead a beginner. Often the people who could benefit most from AI are the ones most likely to be led astray by it, so how we present a recommendation matters as much as whether it’s correct.”

The study’s findings also have implications for medical education and training. As AI becomes more integrated into clinical workflows, training programs may need to teach clinicians how to critically evaluate AI recommendations rather than simply accept them. For non-experts, public health campaigns could emphasize the limitations of AI-based self-diagnosis and the risks of over-reliance on consumer-facing diagnostic tools.

The research was funded by the National Science Foundation, Schmidt Sciences, the National Bureau of Economic Research, and Columbia University. The study authors included MIT graduate student Haoran Zhang, undergraduate Reina Wang, and Luis Soenksen PhD ’20, a research affiliate at the Jameel Clinic, along with clinicians and researchers from multiple institutions.

The central challenge identified by the study—that AI explanations can help experts while misleading beginners—is not unique to dermatology or healthcare. It applies to any domain where AI is used to support decision-making by users with varying levels of expertise, from legal analysis and financial forecasting to engineering design and scientific research. As AI systems become more capable and more widely deployed, the design of their interfaces and explanations will matter as much as the accuracy of their predictions. The one-size-fits-all approach to AI assistance is not just suboptimal; it is actively harmful when it leads users to trust incorrect outputs with unwarranted confidence. The path forward requires not better models alone, but better models designed with a clear understanding of who will use them and how.

Share This Article