Gender and Sexuality Bias in Large Language Models: A Longstanding Issue at an Expanded Scale
-
By
-
Chiara Barbati
-
Virginia Casigliani
-
Caterina Rizzo
-
Anna Odone
-
September 24, 2026
Clinical Scorecard: Gender and Sexuality Bias in Large Language Models: A Longstanding Issue at an Expanded Scale
At a Glance
| Category | Detail |
| Condition | Gender and Sexuality Bias in AI Systems |
| Key Mechanisms | Demographic bias in language generation affecting clinical documentation and reasoning. |
| Target Population | Patients seeking health information from AI systems. |
| Care Setting | Public health concern regarding AI-generated health information. |
Key Highlights
- Gender bias is the most consistently documented form of bias in large language models (LLMs).
- Bias in LLMs can intersect with ethnicity, socioeconomic status, and other attributes.
- Disparities in performance between patient groups are difficult to detect with current accuracy benchmarks.
- Identical cases with different sociodemographic descriptors yield divergent management recommendations.
- Current evaluations often fail to capture the nuances of bias expressed in free text outputs.
Guideline-Based Recommendations
Diagnosis
- Implement sex- and gender-disaggregated evaluation in AI systems.
Management
- Incorporate gender-medicine expertise in benchmark design.
Monitoring & Follow-up
- Stratified monitoring following deployment of AI systems.
Risks
- Potential for biased outputs to exacerbate existing health inequities.
Patient & Prescribing Data
Individuals using AI for health information, particularly those facing cost and access barriers.
Bias in AI outputs may lead to misrepresentation of care needs and treatment allocation.
Clinical Best Practices
- Conduct disaggregated audits to identify subgroup failures in AI outputs.
- Evaluate AI-generated language for bias in word choice, tone, and clinical detail.
Related Resources & Content