Background:
In-hospital heart failure (HF) remains common after primary percutaneous coronary intervention (PPCI) for ST-segment elevation myocardial infarction (STEMI) and is associated with adverse in-hospital outcomes. In addition, whether procedural complete revascularization (CR) can be achieved during the index PCI is clinically relevant but often constrained in real-world practice. We aimed to develop and externally validate machine-learning (ML) models for these two complementary prediction tasks.
Methods:
We conducted a multicenter cohort study of STEMI patients treated with PPCI from three hospitals. Patients from Hezhou People's Hospital (January 2020 to June 2024) comprised the training cohort (n = 734). Patients from two other centers (July 2024 to December 2025) were combined as an independent testing cohort (n = 352). Multiple ML algorithms were benchmarked to predict (1) in-hospital HF and (2) the real-world feasibility of achieving procedural CR during the index PCI. Model performance was assessed using the area under the receiver operating characteristic curve (AUC), the area under the precision-recall curve (AUPRC), classification metrics, calibration curves, decision curve analysis (DCA), and clinical impact curves. Shapley Additive Explanations (SHAP) were used to enhance interpretability.
Results:
For in-hospital HF prediction, CatBoost showed the best overall performance in the independent testing cohort (AUC: 0.973; 95% CI: 0.957–0.989; accuracy: 88.6%), with good calibration and favorable net benefit on DCA. For procedural CR prediction, CatBoost was also selected as the primary model based on its overall performance profile in the independent testing cohort (AUC: 0.970; 95% CI: 0.954–0.987; accuracy: 92.0%), with acceptable calibration and positive net benefit across a broad range of threshold probabilities. Key predictors included LAD involvement, age, symptom-to-guidewire crossing time, and markers related to inflammation, coagulation, renal function, and lipid metabolism.
Conclusions:
In a three-center cohort, we developed and externally validated two ML models for predicting subsequent in-hospital HF after index PPCI and the feasibility of achieving procedural CR during the index PCI. Both models demonstrated good discrimination, calibration, clinical utility, and interpretability, supporting peri-procedural risk stratification and catheterization-laboratory decision support in STEMI patients treated with PPCI.
1 IntroductionAlthough primary percutaneous coronary intervention (PPCI) has become the standard reperfusion strategy for ST-segment elevation myocardial infarction (STEMI), in-hospital heart failure (HF) remains relatively common and is closely associated with adverse in-hospital outcomes (1). Prior studies have shown that HF during the acute phase of STEMI—often reflected by higher Killip class—is not uncommon; moreover, HF present on admission or developing during hospitalization is consistently linked to a higher risk of in-hospital events and mortality, and frequently necessitates higher-acuity monitoring and supportive therapies, thereby increasing healthcare resource utilization (2).
In STEMI patients with multivessel coronary artery disease, complete revascularization (CR) has been demonstrated in randomized trials to reduce ischemia-driven adverse outcomes (3). For example, the COMPLETE trial showed that, compared with culprit-lesion–only PCI, a CR strategy further reduced hard endpoints such as cardiovascular death or myocardial infarction, supporting an overall clinical benefit of CR (4). Accordingly, the 2023 European Society of Cardiology (ESC) guidelines for acute coronary syndromes recommend that, in hemodynamically stable STEMI patients with multivessel disease, non-culprit lesion revascularization may be performed either during the index PCI procedure or within a defined time window, whereas in patients with cardiogenic shock an infarct-related-artery–only approach is generally favored, underscoring the context-dependent nature of CR decision-making (5). However, in real-world practice, whether procedural (index-procedure) CR can be achieved is not determined by anatomy alone. It is jointly influenced by lesion complexity, ischemic burden, procedural risk, hemodynamic status, and operator judgment under time and resource constraints in the catheterization laboratory (6). Therefore, rather than modeling a purely physiological outcome, the present study aimed to develop a prediction model for the real-world likelihood/feasibility of achieving procedural CR during the index PPCI. Such a model may help inform early consideration of immediate vs. staged strategies, facilitate cath-lab workflow and resource planning, and provide a pragmatic basis for identifying patients in whom same-session CR is more or less likely to be achievable (7).
Most prior studies assessing risk in STEMI patients undergoing PPCI—either for (in-hospital) HF or for procedural CR—have been derived from single-center or highly selected cohorts, limiting generalizability and transportability. Even when regression- or machine-learning–based approaches were applied, external validation has often been insufficient and calibration performance incompletely reported, constraining the interpretability of predicted probabilities (8, 9). In addition, evaluation of clinical utility (e.g., decision-curve analysis and net benefit) remains limited, which hampers translation of these models into real-world clinical workflows (10). Recent cardiovascular prediction studies have increasingly incorporated multimodal data, interpretable modeling frameworks, and multicenter validation strategies, underscoring the growing need for clinically usable and transportable risk-prediction tools in real-world settings (11, 12).
In this multicenter study of STEMI patients treated with PPCI, we addressed two complementary prediction tasks. First, we developed and validated a risk model to predict (in-hospital) HF; second, we developed and validated a model to estimate the real-world feasibility/likelihood of achieving procedural CR during the index PCI. We systematically compared multiple ML algorithms in the training cohort and an independent testing cohort in terms of discrimination, calibration, and clinical utility, and applied explainability methods to identify and interpret the key predictors for each endpoint, thereby informing peri-procedural risk stratification and cath-lab decision support.
2 Methods2.1 Study design & data sourceWe conducted a multicenter cohort study using routinely collected clinical data from three hospitals. Patients treated at Hezhou People's Hospital between January 2020 and June 2024 were included as the training cohort for model development. For independent external validation, patients from the other two centers were enrolled between July 2024 and December 2025, and the two datasets were combined to form the validation cohort. The study protocol was approved by the Ethics Committee of Hezhou People's Hospital (approval No. KY2024030605; Supplementary Material 1).
2.2 ParticipantsWe consecutively enrolled STEMI patients admitted to three centers. The inclusion criteria were as follows: (1) meeting the diagnostic criteria of the 2019 guideline for the diagnosis and treatment of acute STEMI (13); and (2) receiving PPCI within 12 h of symptom onset, or undergoing PPCI beyond 12 h when there was clinical and/or electrocardiographic evidence of ongoing/progressive myocardial ischemia. The exclusion criteria were: (1) a history of chronic HF, congenital heart disease, cardiomyopathy, or severe valvular heart disease; (2) incomplete key clinical data; and (3) in-hospital death.
Baseline characteristics and candidate predictors were extracted from routinely collected clinical records, including demographic characteristics, comorbidities, admission laboratory measurements, and selected angiographic/procedural variables. For model development, only variables available up to the predefined prediction time-zero for each endpoint were considered eligible predictors.
2.3 OutcomesTwo relatively independent prediction endpoints were defined: (1) occurrence of (in-hospital) HF during hospitalization; and (2) achievement of CR during the index PCI. Both endpoints were adjudicated based on inpatient medical records, laboratory and imaging data, and catheterization laboratory procedural documentation.
2.3.1 Definition and adjudication window for (in-hospital) HF(In-hospital) HF was defined as the occurrence of clinically recognized acute HF/cardiac dysfunction during the index hospitalization. The endpoint was not based on administrative discharge codes alone, but was adjudicated from the inpatient clinical record using a combination of: (1) clinician-documented new-onset or worsening signs/symptoms consistent with HF or hemodynamic deterioration during hospitalization; (2) supportive imaging or echocardiographic evidence of cardiac dysfunction, when available; and/or (3) initiation or escalation of HF-directed therapies during the hospital stay, such as intravenous diuretics, inotropes, or other circulatory/respiratory supportive measures. The adjudication window spanned from admission to discharge (14, 15). In the present study, this endpoint was treated as a broader in-hospital HF occurrence outcome and was not further subdivided into HF present at admission vs. incident HF developing later during hospitalization.
2.3.2 Definition and adjudication window for procedural CR (index PCI/same-session PPCI)Procedural CR was defined as follows: after completion of the index PCI (same-session PPCI), in addition to treatment of the culprit vessel, any non-culprit lesions with significant stenosis deemed to require revascularization were also treated during the same procedure (e.g., balloon angioplasty and/or stent implantation), thereby achieving complete revascularization within the index procedure (16). Conversely, procedural CR was considered not achieved if, at the end of the index PCI, any significant non-culprit lesions remained untreated (regardless of whether staged PCI was planned during the index hospitalization), or if only culprit-lesion intervention was performed (17).
To reduce temporal ambiguity and minimize the risk of information leakage, prediction time-zero was defined separately for the two endpoints. For the (in-hospital) HF model, prediction time-zero was set at completion of the index PPCI, and the model was designed to predict HF events occurring subsequently during the remainder of hospitalization. Accordingly, eligible predictors for this task were limited to variables available at admission or during the index procedure up to completion of PPCI.
For the procedural CR model, prediction time-zero was defined before final completion of the revascularization strategy during the index PCI. Therefore, only variables available at admission and during procedural assessment up to that time point were considered eligible predictors. Variables directly reflecting subsequent outcome occurrence or clearly downstream management decisions were not intended to serve as proxies for the study outcomes.
2.4 Modeling strategyFor the two relatively independent endpoints—(in-hospital) HF during hospitalization and achievement of CR during the index PCI—we applied a uniform machine-learning workflow while developing and validating two separate models. The dataset was partitioned by center and time period. All candidate predictors retained after dataset curation were entered into the modeling workflow, and no additional outcome-driven univariable screening was performed at the model-building stage. We compared multiple algorithms, including Logistic regression/Lasso, support vector machine (SVM), random forest (RF), gradient boosting decision tree (GBDT), CatBoost, LightGBM, neural networks (NeuralNet), partial least squares (PLS), discriminant analysis, and Bayesian methods. Hyperparameters were tuned within the training cohort using cross-validation in combination with grid search and/or other optimization strategies, as appropriate. Predicted probabilities were converted into binary class labels using a threshold of 0.5. For models implemented within the general training framework, training-cohort performance was summarized mainly on the basis of cross-validation predictions, and these comparative results are additionally reported in the Supplementary Tables to facilitate a more robust assessment of model stability. For CatBoost and LightGBM, training-set predictions were additionally obtained directly from the fitted models; therefore, their apparent training-cohort performance was interpreted cautiously. The analytical dataset was largely complete, with no substantial missingness across the main candidate predictors. Only a small number of variables had occasional missing values, which were handled using a predefined simple imputation strategy according to variable type, with median imputation for continuous variables and mode imputation for categorical variables (18). All analyses were performed in R using the packages caret, pROC, ggplot2, shapviz, kernelshap, dplyr, rmda, catboost, and lightgbm. The main modeling workflow and software environment have been added to improve reproducibility.
In the procedural CR task, Bilateral-or-Multivessel-Disease showed disproportionately high predictive dominance. To reduce over-reliance on a single feature and to assess model robustness, we performed a sensitivity analysis by excluding this variable and repeating model training, hyperparameter tuning, and evaluation using the same workflow.
The independent testing cohort was used as the principal benchmark for model comparison and interpretation. Given the possibility of optimistic performance estimates in the training cohort, the primary analyses emphasized discrimination, calibration, and clinical utility in the independent testing cohort rather than the apparent performance in the training cohort. For the HF prediction task, the final interpretation of model behavior relied primarily on the dominant features identified by SHAP, which were mainly demographic, laboratory, ischemic-time, and angiographic/procedural variables rather than medication-related variables.
2.5 Performance evaluationModel discrimination was primarily assessed using the area under the receiver operating characteristic curve (AUROC), with the area under the precision–recall curve (AUPRC) additionally reported. Classification performance metrics included accuracy, F1-score, positive predictive value (PPV), negative predictive value (NPV), sensitivity, specificity, and the Youden index. Calibration was evaluated using calibration curves. Clinical utility was assessed using decision curve analysis (DCA) and clinical impact curves to quantify net benefit across a range of threshold probabilities and to illustrate potential clinical value.
2.6 ExplainabilityFor the best-performing models, we used Shapley Additive Explanations (SHAP) to enhance interpretability. Global feature importance and SHAP beeswarm plots were used to identify key predictors and their directional contributions, while SHAP dependence plots were generated to depict the relationships between pivotal features and model outputs. In addition, individual-level explanations were provided to support clinical interpretability and facilitate potential implementation (19).
3 Results3.1 Patient selection and cohort characteristicsThe patient selection process is shown in Figure 1. After application of the predefined inclusion and exclusion criteria, a total of 1,086 eligible STEMI patients treated with PPCI were included in the final analysis, comprising 734 patients in the training cohort and 352 patients in the independent testing cohort. Baseline characteristics were well balanced between the two cohorts (Table 1). The two cohorts showed similar distributions of demographic features (e.g., age and gender), comorbidities, laboratory indices, and other clinical/procedural variables, with no statistically significant differences observed across baseline variables (all P > 0.05). These findings support the comparability of the cohorts and the suitability of the testing cohort for independent validation. The same training/testing cohort partition was applied to both prediction tasks, namely (in-hospital) HF prediction and procedural CR feasibility prediction.

Flowchart showing the screening and selection of STEMI patients treated with PPCI from three participating centers, including application of the predefined inclusion and exclusion criteria and assignment to the training cohort (n = 734) and independent testing cohort (n = 352). Both endpoints—(in-hospital) HF prediction and procedural CR feasibility prediction—were modeled using the same training/testing framework.
VariablesTrainingBaseline characteristics and outcome-related differences of all variables in the training and testing cohorts.
Bold values indicate statistical significance (P < 0.05).
Baseline demographic, clinical, laboratory, and procedural variables were compared between the training cohort (Hezhou People's Hospital, January 2020–June 2024; n = 734) and the independent testing cohort (two external centers, July 2024–December 2025; n = 352). Continuous variables are presented as mean (standard deviation) or median (interquartile range), as appropriate; categorical variables are presented as n (%). P values were calculated using appropriate statistical tests according to data type and distribution.
COPD, chronic obstructive pulmonary disease; MI, myocardial infarction; CKD, chronic kidney disease; WBC, white blood cell count; NEUT, neutrophil count; Hb, hemoglobin; PLT, platelet count; hsCRP, high-sensitivity C-reactive protein; HDL-C, high-density lipoprotein cholesterol; LDL-C, low-density lipoprotein cholesterol; ALB, albumin; UA, uric acid; BUN, blood urea nitrogen; LAD, left anterior descending artery; LCX, left circumflex artery; RCA, right coronary artery; ACEI, angiotensin-converting enzyme inhibitor; ARB, angiotensin receptor blocker; ARNI, angiotensin receptor–neprilysin inhibitor.
3.2 Outcome 1: in-hospital HF prediction model3.2.1 Comparison of HF and Non-HF baselinesAcross both the training cohort (Table 2) and the testing cohort (Table 3), baseline characteristics showed a consistent pattern of differences between the HF and non-HF groups. Compared with the non-HF group, patients who developed HF were older, more frequently had LAD-related lesions, and had a longer symptom-to-guidewire crossing time (Time to Guidewire) (both cohorts, P < 0.001), indicating that prolonged ischemic time and delayed reperfusion may contribute to a higher risk of (in-hospital) HF. In terms of coronary anatomy, the HF group had a higher prevalence of Bilateral or Multivessel Disease (both cohorts, P < 0.001), further indicating a greater overall atherosclerotic burden.
VariablesHF (n = 307)Non-HF (n = 427)P-valueDemographics Gender (male, %)198 (64.5%)272 (63.7%)0.824 Age (years)69.12 ± 9.4560.23 ± 10.12<0.001 Smoking (%)78 (25.4%)105 (24.6%)0.798Comorbidities Diabetes (%)115 (37.5%)82 (19.2%)<0.001 Hypertension (%)205 (66.8%)195 (45.7%)<0.001 COPD (%)22 (7.2%)14 (3.3%)0.015 Atrial Fibrillation (%)28 (9.1%)12 (2.8%)<0.001 Malignant Neoplasms (%)6 (2.0%)8 (1.9%)0.945 Previous MI (%)35 (11.4%)22 (5.2%)0.002 Previous Coronary Revasc (%)22 (7.2%)18 (4.2%)0.078 CKD (%)38 (12.4%)15 (3.5%)<0.001Laboratory WBC (109/L)9.24 ± 2.856.45 ± 1.92<0.001 NEUT (109/L)6.52 ± 2.154.18 ± 1.45<0.001 Hb (g/L)122.5 ± 18.4138.2 ± 14.2<0.001 PLT (109/L)215 ± 68212 ± 620.542 hsCRP (mg/L)9.85 (4.21, 22.45)1.52 (0.65, 3.24)<0.001 Cholesterol (mmol/L)4.28 ± 1.124.15 ± 0.880.082 Triglycerides (mmol/L)1.82 (1.25, 2.12)0.012 HDL-C (mmol/L)1.02 ± 0.211.18 ± 0.28<0.001 LDL-C (mmol/L)2.88 ± 0.942.42 ± 0.75<0.001 ALB (g/L)36.52 ± 5.1242.48 ± 3.65<0.001 UA (umol/L)412 ± 95328 ± 72<0.001 BUN (mmol/L)8.82 (5.45, 13.12)4.55 (3.52, 6.28)<0.001 D-Dimer (mg/L)1.45 (0.68, 3.25)0.22 (0.11, 0.42)<0.001Angiographic Time to Guidewire (h)8.95 ± 3.424.12 ± 1.85<0.001 LAD Occlusion (%)188 (61.2%)155 (36.3%)<0.001 LCX Occlusion (%)82 (26.7%)108 (25.3%)0.672 RCA Occlusion (%)88 (28.7%)112 (26.2%)0.456 Bilateral/Multivessel (%)142 (46.3%)95 (22.2%)<0.001 Overall complete revascularization during index hospitalization (%)145 (47.2%)265 (62.1%)<0.001Procedural Postoperative Slow Flow (%)45 (14.7%)12 (2.8%)<0.001 Postoperative No Reflow (%)15 (4.9%)5 (1.2%)0.002Medication ACEI ARB ARNI (%)135 (44.0%)182 (42.6%)0.715 Beta blockers (%)192 (62.5%)255 (59.7%)0.442Comparison of baseline characteristics between HF and Non-HF patients in the training cohort.
Bold values indicate statistical significance (P < 0.05).
Baseline variables were compared between patients who developed (in-hospital) HF and those without (in-hospital) HF in the training cohort (n = 734). Data presentation and statistical testing are as described in Table 1.
HF, heart failure; LAD, left anterior descending artery; Time-to-Guidewire, symptom-to-guidewire crossing time; other abbreviations as in Table 1.
VariablesHF (n = 171)Non-HF (n = 181)P-valueDemographics Gender (male, %)108 (63.2%)110 (60.8%)0.648 Age (years)69.45 ± 8.9260.84 ± 9.85<0.001 Smoking (%)43 (25.1%)45 (24.9%)0.952Comorbidities Diabetes (%)64 (37.4%)41 (22.7%)<0.001 Hypertension (%)115 (67.3%)82 (45.3%)<0.001 COPD (%)12 (7.0%)7 (3.9%)0.185 Atrial Fibrillation (%)15 (8.8%)5 (2.8%)0.012 Malignant Neoplasms (%)3 (1.8%)4 (2.2%)0.752 Previous MI (%)19 (11.1%)9 (5.0%)0.035 Previous Coronary Revasc (%)11 (6.4%)10 (5.5%)0.724 CKD (%)21 (12.3%)4 (2.2%)<0.001Laboratory WBC (109/L)9.32 ± 2.626.
Comments (0)