91°µÍř

Can AI help make medical records less biased? New study suggests yes—with caveats

Body

Large language models can identify judgmental language in clinical notes but the settings play a major role in accuracy. 

Teenu Xavier PhD, RN. Photo provided

“Addict,” “non-compliant,” “failed treatment,” and “obese person” are examples of stigmatizing language that can appear in medical records. At 91°µÍř’s College of Public Health, researchers are exploring whether artificial intelligence (AI) can help identify this kind of language in clinical notes before it impacts patient care.

Nurse scientist and colleagues found that large language models (LLMs) show promise in identifying stigmatizing language in clinical documentation, but their performance is highly dependent on their settings. Model size, temperature settings, prompting strategies, and even note type can substantially influence results.

One finding was consistent across every model tested: Providing examples of stigmatizing language improved accuracy.

“Simply selecting an LLM is not enough when used for clinical documentation,” said Xavier, an assistant professor in the School of Nursing. “Careful attention must be paid to settings and prompting before these tools can be reliably used in health care environments.”

Why does this matter?

The use of stigmatizing language in clinical documentation can reinforce bias and affect a patient’s future care. AI tools may be able to help identify this kind of language, promoting more equitable care and improving patient trust and experience.

“Pre-trained models, when optimized for identifying stigmatizing language, could help enable more timely interventions and modifications to the documentation process,” Xavier said. “Our research highlights the need for continued collaboration between health care professionals and AI developers to create tools that improve communication, reduce bias, and improve the overall patient experience.”

What are the detailed study findings?

  • The largest LLM (trained on large amounts of data) was the best at predicting “stigmatizing” language (94%), but the worst at predicting “not stigmatizing” correctly (47%).
     
  • The smallest LLM was the best at predicting “not stigmatizing” correctly (99.7%), but worst at predicting “stigmatizing” correctly (2%).
     
  • When researchers gave the LLM an example of stigmatizing language, accuracy improved in all models.
     
  • Emergency provider notes were most accurately (69%) categorized as “stigmatizing” or “non-stigmatizing,” and plan of care notes had the lowest accuracy (56%). Misclassifications most commonly arose in long, clinically dense notes where neutral descriptions of complex illness or adverse events were mistaken by the models for judgmental language.
     
  • Larger models worked best at lower temperature (how predictable or random the LLMs’s output is when making a classification) and smaller models improved with higher temperature, which means the LLM took more risks in interpretation.

was published in JAMIA Open in April 2026. Co-authors include Jane M. Carrington from the University of Florida and Joshua Lambert from the University of Cincinnati.

Thumbnail photo by from Adobe Stock.

Key Takeaways

  • A George Mason study found that large language models show promise in identifying stigmatizing language in clinical documentation, but their performance is highly dependent on their settings, such as model size, temperature settings, prompting strategies, and note type. 
     
  • The ability to detect and correct stigmatizing language early can reduce bias in patient care and lead to improved patient trust and better health outcomes. 
     
  • Continued collaboration between health care workers and AI developers is needed to create tools that enhance communication, reduce bias, and improve the overall patient experience.