Clinical Report: Utilizing Explainable Machine Learning to Classify Hypertension Prevalence
Overview
This study employed explainable machine learning models to classify hypertension prevalence among 4,606 residents in Hainan Province, China. The eXtreme Gradient Boosting (XGBoost) model demonstrated the highest performance with an AUC of 0.8461.
Background
Hypertension is a major public health concern, significantly contributing to cardiovascular disease and mortality worldwide. The prevalence of hypertension has been rising, particularly in China, where awareness and management remain low.
Data Highlights
Metric
Value
Sample Size
4,606
Hypertension Prevalence
32.5%
Training Set Size
3,224 (70%)
Validation Set Size
1,382 (30%)
AUC of XGBoost Model
0.8461
Key Findings
The study included 4,606 residents, with a hypertension prevalence of 32.5%.
The random forest algorithm identified the top 10 features influencing hypertension classification.
The XGBoost model achieved the highest classification performance with an AUC of 0.8461.
SHAP values were utilized to explain feature contributions to the model.
The LIME algorithm provided individual classification explanations, enhancing model interpretability.
Clinical Implications
The findings indicate that machine learning models, particularly XGBoost, can classify hypertension prevalence.
Conclusion
The development of an explainable machine learning model for hypertension classification demonstrates its potential for screening.