To evaluate the effectiveness of the aiDIVA hybrid artificial intelligence system in ranking disease-causing variants for rare disease cases.
Approach:
System Development: aiDIVA combines evidence-based scoring, machine learning, phenotype matching, inheritance information, and large language models to generate a ranked list for clinical assessment.
Training Data: The model was trained using over 155,000 variants from ClinVar, including approximately 90,000 pathogenic and 65,000 benign variants.
Evaluation Methodology: The system was benchmarked against 3,041 rare disease cases previously solved by genetics specialists, with separate models for dominant and recessive inheritance.
Key Findings:
aiDIVA placed the causal variant within the top three in 97% of cases and at rank one in 88% among cases with variants in ClinVar or the Human Gene Mutation Database.
In an independent group of 1,014 cases, aiDIVA placed 93% of causal variants within the top three and 96% within the top 10.
aiDIVA identified 45 previously unsolved cases as newly solved after expert review.
Interpretation:
aiDIVA is a decision-support tool that aids in variant triage and periodic reanalysis, but final classification requires expert review.
Limitations:
Language-model responses were inconsistent in a small proportion of tests and could cite nonexistent or irrelevant literature.
Sending phenotype data to cloud-based models raises privacy considerations.
aiDIVA does not currently support structural variants and has lower performance for in-frame insertions and deletions compared to single-nucleotide variants.
Conclusion:
aiDIVA demonstrates high accuracy in ranking disease-causing variants, supporting its potential use in clinical settings.