TY - JOUR T1 - Explainable Ensemble Machine Learning Framework for Lung Cancer Status Classification: A Retrospective Cross-sectional Study AU - Saad, Hamza JF - Cancer Screening and Prevention VL - IS - 000 SN - 2835-3315 SP - EP - Y1 - 2026-09-30 DO - 10.14218/CSP.2026.00009 UR - https://www.xiahepublishing.com/2835-3315/CSP-2026-00009 AB - Background and objective Machine-learning approaches that combine predictive performance with model interpretability may improve lung cancer status classification. This study aimed to develop and internally evaluate an Explainable Precision Screening Framework for lung cancer risk classification using demographic, behavioral, and symptom-based variables from a publicly available dataset. Methods This retrospective cross-sectional study analyzed a publicly available Kaggle dataset containing 309 records, including 270 labeled as lung cancer and 39 as non-cancer. Six models—logistic regression, support vector machine (SVM), random forest, LightGBM, XGBoost, and a stacking ensemble—were compared using a stratified hold-out test set and repeated stratified five-fold cross-validation. Performance was assessed using classification metrics with bootstrap 95% confidence intervals (CIs). A separate Shapley additive explanations (SHAP) analysis was applied to the standalone SVM model for exploratory feature attribution. Results Across all models, accuracy ranged from 85% to 92% (ROC-AUC: 0.93–0.95). The stacking ensemble achieved 0.92 accuracy (95% CI: 0.85–0.98), 0.94 sensitivity (95% CI: 0.88–1.00), 0.75 specificity (95% CI: 0.40–1.00), 0.96 precision (95% CI: 0.89–1.00), 0.95 F1 score (95% CI: 0.91–0.99), and 0.95 ROC-AUC (95% CI: 0.89–0.99). Its PR-AUC was 0.993 (95% CI: 0.981–0.999), and its Brier score was 0.074 (95% CI: 0.037–0.121). SHAP analysis identified smoking, yellow fingers, coughing, chest pain, wheezing, shortness of breath, and age as the features contributing most strongly to its predictions. Conclusions Within this public retrospective dataset, the stacking ensemble was among the highest-performing models, whereas separate SHAP analysis of the standalone SVM model identified the features contributing most strongly to its predictions. Given the marked class imbalance, the absence of a reported clinical reference standard for the source labels, the framework requires independent validation in clinically verified multicenter cohorts before clinical use.