| Cardiovascular disease remains a leading global health challenge, highlighting the need for reliable and clinically aligned risk prediction tools. This study evaluates a range of classical and gradient boosting machine learning models using a large dataset of 70,000 structured clinical records. Rather than relying solely on default-threshold benchmarking, we adopt a decision-aware evaluation strategy in which classification thresholds are selected under a predefined sensitivity constraint (Sensitivity ≥ 0.90) to reflect screening priorities. Model training and assessment are repeated across multiple stratified splits to quantify performance variability and reduce split-dependent bias. Gradient boosting models achieved the highest discrimination while maintaining stable operating characteristics under the high-sensitivity policy. Probability calibration was assessed using Brier scores and reliability curves, indicating reliable risk estimates, particularly for LightGBM. SHAP-based analysis demonstrated consistent reliance on established cardiovascular risk factors, and examination of false-negative cases provided additional insight into model limitations. By integrating threshold-aware evaluation, repeated experimentation, and explanation stability analysis, this paper provides a deployment-aware assessment framework for interpretable cardiovascular risk screening. |
*** Title, author list and abstract as submitted during Camera-Ready version delivery. Small changes that may have occurred during processing by Springer may not appear in this window.