Sex and gender bias in large language models: an old problem at a new scale - Summary - MDSpire
Coming Soon: Introducing MDSpire News. Learn more
Conexiant’s news site is now MDSpire News. Learn more

Gender and Sexuality Bias in Large Language Models: A Longstanding Issue at an Expanded Scale

  • By

  • Chiara Barbati

  • Virginia Casigliani

  • Caterina Rizzo

  • Anna Odone

  • September 24, 2026

Share

Objective:

To address the issue of gender bias in large language models (LLMs) and its implications for healthcare, emphasizing the need for sex- and gender-disaggregated evaluations.

Approach:
  • Evidence Review: The article reviews existing literature on gender bias in LLMs, highlighting its prevalence and impact on clinical documentation and reasoning.
  • Benchmark Design: It advocates for the incorporation of gender-medicine expertise in benchmark design and stratified monitoring of LLM outputs.
Key Findings:
  • Gender bias is the most consistently documented form of bias in LLMs, appearing in 15 of 16 studies reviewed.
  • Bias in LLM outputs can exacerbate existing health inequities, particularly for disadvantaged groups.
  • Clinical implications of gender bias include differential treatment recommendations and diagnostic rankings based on patient demographics, such as the use of complex wording for male patients and euphemistic language for female patients.
  • Bias is bidirectional and context-dependent, affecting both male and female patients in varying ways.
Interpretation:

The evidence indicates that LLMs generate biased outputs, but it remains unclear how these biases affect clinical decisions and patient outcomes in real-world settings.

Limitations:
  • Most evaluations focus on single protected attributes, missing intersectional harms.
  • Current metrics fail to capture biases expressed through language in generative models.
  • The evidence primarily comes from controlled evaluations, lacking prospective studies involving actual patients.
Conclusion:

The incorporation of gender equity as a design requirement in LLMs is still unresolved, necessitating further research and evaluation.

Sources:

Original Source(s)

Related Content