Sex and gender bias in large language models: an old problem at a new scale - Report - MDSpire
Coming Soon: Introducing MDSpire News. Learn more
Conexiant’s news site is now MDSpire News. Learn more

Gender and Sexuality Bias in Large Language Models: A Longstanding Issue at an Expanded Scale

  • By

  • Chiara Barbati

  • Virginia Casigliani

  • Caterina Rizzo

  • Anna Odone

  • September 24, 2026

Share

Clinical Report: Gender and Sexuality Bias in Large Language Models

Background

As patients increasingly rely on LLMs for health information, the potential for demographic bias in these systems raises significant public health concerns. Gender bias has been consistently documented and intersects with other demographic factors, potentially exacerbating existing health inequities.

Data Highlights

No numerical or trial data was provided in the source material.

Key Findings

  • Gender bias is the most consistently documented form of bias in LLMs, appearing in 15 of 16 studies.
  • LLMs can exhibit various biases, including racial, ethnic, age, and socioeconomic biases, alongside gender bias.
  • Clinical documentation may reflect gender bias, with male patients receiving more complex wording compared to female patients.
  • OpenAI GPT-4 has been shown to perpetuate demographic stereotypes in clinical reasoning tasks.
  • Bias in LLM outputs can lead to divergent management recommendations based solely on sociodemographic descriptors.
  • Gender bias in LLMs is bidirectional and context-dependent, affecting both men and women.

Clinical Implications

Healthcare professionals should be aware of the potential for biased outputs from LLMs, which may impact clinical decision-making and patient care.

Conclusion

Addressing gender and sexuality bias in LLMs is critical.

Related Resources & Content

  1. Journal of Medical Internet Research, 2026 -- Evaluating the Potential of Reasoning Large Language Models to Perpetuate Racial and Gender Disease Stereotypes in Health Care
  2. npj Digital Medicine, 2025 -- The evaluation illusion of large language models in medicine
  3. Journal of Medical Internet Research, 2026 -- Health Care Professionals' Perspectives on the Integration and Regulation of Large Language Models: A Cross-Sectional Survey Analysis
  4. Frontiers in Psychiatry — Gender differences in autism prevalence: origins of bias and its current scientific relevance
  5. HTI-1 Final Rule - ONC - Office of the National Coordinator for Health Information Technology
  6. Considerations for Generative AI in Public Health | Artificial Intelligence | CDC
  7. Consensus framework for the validation of generative AI: call for collaborators on the Validation Accords | Nature Medicine
  8. AMA policies to ensure AI supports—not replaces—physician judgment | American Medical Association
  9. Bias Evaluation in Medical Applications of Large Language Models: A Systematic Review - PMC
  10. Large language models exhibit stigmatizing behaviour in contextual judgements of health conditions | Nature Health
  11. Large language models provide unsafe answers to patient-posed medical questions - PMC
  12. Journal of Medical Internet Research - Evaluating the Potential of Reasoning Large Language Models to Perpetuate Racial and Gender Disease Stereotypes in Health Care
  13. Evaluating anti-LGBTQIA+ medical bias in large language models | PLOS Digital Health
  14. Auditing Large Language Model–Generated Digital Standardized Patients for Demographic Bias: A Simulation Study with HIV Pre-Exposure Prophylaxis Screening as a Tracer Condition | medRxiv
  15. Measuring stereotype and deviation biases in large language models | Scientific Reports
  16. Mitigating Automation Bias in Physician-LLM Diagnostic Reasoning Using Behavioral Nudges: A Randomized Controlled Trial | medRxiv
  17. Evidence, use cases, and implementation safeguards of large language models in primary care | Communications Medicine

Original Source(s)

Related Content