Sex and gender bias in large language models: an old problem at a new scale - Scorecard - MDSpire
Coming Soon: Introducing MDSpire News. Learn more
Conexiant’s news site is now MDSpire News. Learn more

Gender and Sexuality Bias in Large Language Models: A Longstanding Issue at an Expanded Scale

  • By

  • Chiara Barbati

  • Virginia Casigliani

  • Caterina Rizzo

  • Anna Odone

  • September 24, 2026

Share

Clinical Scorecard: Gender and Sexuality Bias in Large Language Models: A Longstanding Issue at an Expanded Scale

At a Glance

CategoryDetail
ConditionGender and Sexuality Bias in AI Systems
Key MechanismsDemographic bias in language generation affecting clinical documentation and reasoning.
Target PopulationPatients seeking health information from AI systems.
Care SettingPublic health concern regarding AI-generated health information.

Key Highlights

  • Gender bias is the most consistently documented form of bias in large language models (LLMs).
  • Bias in LLMs can intersect with ethnicity, socioeconomic status, and other attributes.
  • Disparities in performance between patient groups are difficult to detect with current accuracy benchmarks.
  • Identical cases with different sociodemographic descriptors yield divergent management recommendations.
  • Current evaluations often fail to capture the nuances of bias expressed in free text outputs.

Guideline-Based Recommendations

Diagnosis

  • Implement sex- and gender-disaggregated evaluation in AI systems.

Management

  • Incorporate gender-medicine expertise in benchmark design.

Monitoring & Follow-up

  • Stratified monitoring following deployment of AI systems.

Risks

  • Potential for biased outputs to exacerbate existing health inequities.

Patient & Prescribing Data

Individuals using AI for health information, particularly those facing cost and access barriers.

Bias in AI outputs may lead to misrepresentation of care needs and treatment allocation.

Clinical Best Practices

  • Conduct disaggregated audits to identify subgroup failures in AI outputs.
  • Evaluate AI-generated language for bias in word choice, tone, and clinical detail.

Related Resources & Content

Original Source(s)

Related Content