Development and Validation of an Interpretable Machine Learning Model for Prediction of the Need for Surgical Evacuation in Patients with Incomplete Abortion

Introduction

Incomplete abortion refers to partial expulsion of gestational products with retained products of conception (RPOC) remaining in the uterine cavity.1 As a common reproductive public health issue among women of childbearing age, its main manifestations are persistent and heavy vaginal bleeding. Without timely intervention, it may cause hypovolemic shock,2 uterine hematoma, and lead to intrauterine infection, endometritis, intrauterine adhesions, secondary infertility,3 as well as psychological distress such as anxiety,4 imposing substantial health burdens and unnecessary medical costs on individuals and healthcare systems. Targeted risk stratification may theoretically mitigate complications, preserve fertility and rationalize resource allocation to reduce heterogeneity in clinical management. Yet no standardized quantitative stratification tools are available for incomplete abortion, leading to inconsistent clinical practice across providers and facilities.

The initial treatment for incomplete abortion typically involves drug regimens such as mifepristone and misoprostol.5 However, choosing between surgical and conservative follow-up therapy remains a key population-level challenge in reproductive public health: curettage rapidly removes residual tissue but may induce endometrial injury;6 although hysteroscopic electrosurgery improves surgical precision, it still carries a risk of thermal injury.3,7 Current decisions are primarily based on clinicians’ professional judgment, with no unified quantitative thresholds to assist with surgical intervention risk assessment at the population health level. This may result in overtreatment,8 which reduces patient satisfaction and increases the risk of medical disputes. Therefore, identifying key predictive indicators, exploring their quantifiable risk thresholds, and developing promising models may potentially help improve public health management and refine clinical practice.

In recent years, artificial intelligence (AI) has opened new avenues for the prediction of obstetric and gynecological disorders.9 As a multifactorial public health issue, incomplete abortion involves complex interactive pathogenic factors,10–13 and AI has been proven to integrate multi-dimensional clinical data to identify potential patterns for predicting treatment outcomes of early pregnancy loss.14 However, the interpretability of current AI remains controversial, which limits clinicians’ ability to evaluate AI models.15 In addition, most studies rely on conventional clinical indicators or single machine learning models with poor generalization.16,17 Although interpretable AI tools such as SHAP and the integration of multi-dimensional clinical data have been widely used in predictive modeling and applied to the field of incomplete abortion,18,19 few studies have systematically combined these methods with ensemble modeling and used interpretable approaches to quantify clinical risk thresholds, indicating the necessity of conducting targeted research to address the clinical heterogeneity of incomplete abortion.20 Therefore, integrating artificial intelligence with interpretability analysis to assist clinical decision-making for incomplete abortion is not only a technological innovation but also an urgent need to address this key public health problem.

Accordingly, this study constructed a stacking ensemble framework that integrates multi-dimensional data including clinical symptoms,21 ultrasound indicators and biochemical parameters to overcome the performance limitations of single models. We used explainable analyses to quantify the impacts of key features on surgical intervention risk and provide clinical thresholds for personalized risk stratification. This study provides preliminary supporting evidence for individualized clinical decision-making, and puts forward exploratory quantitative criteria. These criteria may act as a reference for future research on clinical decision guidance and reproductive health policy development only after rigorous multi-centre external validation.

Materials and Methods Research Framework

This study was divided into five phases, as illustrated in Figure 1. First, clinical data were retrospectively collected and preprocessed; second, a stacking ensemble framework was constructed using random forest (RF), gradient boosting decision tree (GBDT), extreme gradient boosting (XGBoost), and light gradient boosting machine (LightGBM) as base learners and logistic regression (LR) as the meta-learner; third, model performance was comprehensively evaluated via internal validation and hold-out validation with multiple performance metrics; fourth, feature importance was assessed using a weighted average of IG and SHAP values, followed by visualization analyses to enhance model interpretability; and finally, critical risk thresholds for key risk factors were quantified using dependence plots.

A flowchart detailing a clinical study process from data collection to decision support.

Figure 1 Research framework flowchart.

Study Population and Inclusion/Exclusion Criteria

This study was conducted in accordance with the Declaration of Helsinki and relevant ethical guidelines, and was approved by the Ethics Committee of Hangzhou Women’s Hospital (Approval Number: 2021‑A‑04) on 2 June 2021. It retrospectively selected 411 patients who voluntarily requested medical abortion at Hangzhou Obstetrics and Gynecology Hospital between 2023 and 2025 as the study population. These patients were divided into two groups based on subsequent treatment: the surgical treatment group (111 patients) and the non-surgical treatment group (300 patients). The inclusion and exclusion criteria are as follows:

Inclusion Criteria: Women aged 20–40 years; Singleton pregnancies, amenorrhea ≤ 70 days, voluntarily choosing medical abortion to terminate early pregnancy; No contraindications to mifepristone or prostaglandin medications; Positive urine or β-hCG; Ultrasound examination showing no complete gestational sac in the uterine cavity, replaced by strong echogenic or mixed echogenic signals with or without peripheral blood flow.

Exclusion Criteria: Suspected ectopic pregnancy or conditions where trophoblastic disease cannot be excluded; Long-term systemic corticosteroid treatment, or patients with chronic adrenal insufficiency, coagulation disorders, or those currently receiving anticoagulant therapy; Uterine malformations; Ultrasound follow-up within 3 days after medical abortion showing retained pregnancy tissue requiring immediate curettage, or massive bleeding requiring curettage; Missed abortion patients; Patients with incomplete clinical data or those unwilling to participate in follow-up.

Outcome Definition

The primary outcome was defined as surgical treatment for incomplete abortion after medical abortion.

Surgical treatment group: Patients required surgery due to ultrasound-confirmed RPOC, persistent heavy vaginal bleeding (> 2 weeks), severe abdominal pain, or intrauterine infection. Surgery was also indicated if RPOC and elevated β-hCG persisted despite 1–2 cycles of conservative treatment, or if urgent surgery was required for heavy bleeding or severe pain after menstrual resumption.

Non-surgical treatment group: Patients showed complete expulsion of RPOC on ultrasound, with significantly decreased β-hCG, resolved vaginal bleeding within 2 weeks, and no infection or severe symptoms. No further surgical intervention was needed.

Collection of Influencing Factors

The data collected in this study include quantitative and qualitative variables from 411 patients. The indicators cover basic patient information, preoperative evaluation, postoperative monitoring, and outcome variables. Core data are presented in Table 1. This study defined three key time points related to medical abortion:

Table 1 Core Data Collection

Pre-medical abortion examination: positive medical history (PMH), gestational sac diameter on pre-medical abortion ultrasound (GSD-PMAUS), days post last menstrual period (DPLMP), medication history (MH), Uterine Position (UP), etc.

First post-medical abortion follow-up (7–14 days): residual tissue size at first post-medical abortion review (RTS-FPAR), β-hCG, RI, post-medical abortion complications, etc.

Second follow-up (post-resumption of menstruation): ultrasound-detected residual tissue, menstruation resumption time (MRT), etc.

Data Preprocessing

Data preprocessing was performed to address errors arising from data entry, extraction, transformation, and loading. The procedures included missing value imputation, outlier removal, and data standardization.22–24 Quantitative variables (e.g, high-density lipoprotein [HDL]) were imputed by mean values to maintain their original data distribution. Qualitative variables (e.g, number of previous abortions [NPA]) were filled with mode values to reduce bias during model training. Outliers were detected and removed using the Rajda criterion at a 99.7% confidence level, as shown in Equation (1):

(1)

Construction of Predictive Models Four Common Ensemble Learning Models

Ensemble learning models perform well in high‑dimensional and noisy clinical data with good generalization and robustness.25 We therefore selected four typical tree-based ensemble models as base learners: RF, GBDT, XGBoost, and LightGBM.

RF generates many independent decision trees, which helps reduce overfitting and improve model stability.26 GBDT optimizes residuals iteratively and performs well in capturing nonlinear relationships in clinical data.27 XGBoost uses regularization and second-order derivatives to improve prediction accuracy and avoid overfitting.28 LightGBM adopts histogram-based feature discretization to speed up training and handle high-dimensional data efficiently.29 For comparison, this study also developed several conventional machine learning models, including logistic regression (LR), support vector classification (SVC), and k-nearest neighbors (KNN).

Stacking Framework

This study employed a two-layer stacking ensemble framework with base learners and a meta-learner.21 This strategy preserves the unique strengths of each base model, mitigates overfitting, and performs favorably in complex clinical prediction tasks involving multiple interacting factors, making it ideal for our core aim: predicting the need for surgical intervention after medical abortion.

Accordingly, this study built a two-layer stacking model. The first layer included four complementary base learners, designed to capture complex data patterns. The second layer employed LR as the meta-learner for its low complexity and high clinical interpretability. An equal-weight strategy was used to integrate outputs from base learners, and stratified 5-fold cross-validation was applied throughout modeling to avoid data leakage and alleviate class imbalance. The implementation steps were as follows:

First, the preprocessed dataset was stratified into 5 folds according to outcome distribution. Each fold was used as the validation set in turn, while the remaining 4 folds formed the training set.

Second, all base learners were trained separately on the training sets, and their predicted probabilities of surgical evacuation were aggregated to construct a new 4-dimensional feature matrix for the meta-learner.

Last, the LR meta-learner was trained on the new feature matrix to generate the final prediction of surgical evacuation requirement.

The schematic diagram of the stacking method is shown in Figure 2.

A diagram of a stacking ensemble model with training and testing processes.

Figure 2 Stacking ensemble model architecture diagram.

Hyperparameter Tuning

All hyperparameter tuning was performed using stratified 5‑fold cross‑validation to preserve the class ratio in each fold and ensure unbiased hyperparameter selection. Grid search was used to optimize key hyperparameters for LR and SVC; hyperparameters for KNN, RF, GBDT, XGBoost, and LightGBM were tuned manually. This study uses predicted probabilities from each base learner as new input features for the meta-learner in the stacking framework.

Model Evaluation

This study developed a multi-dimensional evaluation system to assess model discrimination, stability, and generalization ability, ensuring objective and reliable results. This system included two key parts: a validation design and standardized performance metrics.

Model Validation

Internal Validation: This study used stratified 5-fold cross-validation to reduce overfitting and obtain stable results. The internal dataset included 311 cases (84 surgical, 227 non-surgical), split into five subsets with matching class ratios. In each fold, one subset was used for validation and the rest for training. Grid search optimized hyperparameters, and we reported the average performance across all five folds.

Hold‑out Validation: A separate hold-out subset was used to evaluate the model’s generalization and clinical applicability. This subset included 100 cases (27 surgical, 73 non-surgical). All protocols and definitions were consistent with the internal validation dataset, including inclusion/exclusion criteria, feature selection, data collection, follow-up, and outcome definition. The final model was applied directly without further adjustment, and its predictive performance was evaluated.

Performance Evaluation Metrics

This study assessed model performance in internal and hold-out validation using four clinical metrics: area under the curve (AUC), accuracy, sensitivity, and specificity. AUC was the main measure of discrimination and all results include 95% CIs. Formulas for accuracy, sensitivity, and specificity appear in Equations (2)–(4):

(2)

(3)

(4)

where TP = true positive, FP = false positive, TN = true negative, and FN = false negative.

Feature Importance Analysis and SHAP Visualization

We used IG and SHAP values to assess feature importance and improve the model’s clinical interpretability. Through this analysis, we identified the key features that most strongly influence model predictions of surgical treatment need, clarified core indicators for the model, and laid a foundation for developing risk-based clinical management tools to optimize personalized care for patients with incomplete abortion. Additionally, SHAP visualization converted abstract relationships between features and risks into interpretable graphical evidence, further strengthening the clinical applicability of the model.

Feature Importance Analysis Based on IG

IG quantifies feature importance by accumulating IG at each split node in the decision tree ensemble. The global importance of feature () is the average importance across all decision trees (Equation (5)):

(5)

where is the number of decision trees, represents the -th decision tree, and represents the importance of feature in .

In a single tree, feature importance is the sum of IG values across all non-leaf nodes split by feature (Equation (6)):

(6)

where is the IG of feature at node , is an indicator function (1 if feature is used for splitting, 0 otherwise), and is the number of non-leaf nodes.

SHAP‑Based Feature Importance and Visualization Analysis

Based on cooperative game theory, SHAP values measure the marginal contribution of each feature to model prediction.30 The SHAP value of feature is calculated as (Equation (7)):

(7)

where is the SHAP value, is a subset of features that does not include , and are the numbers of features in and the full set , and is the model prediction using subset .

SHAP visualizations were used to explain feature importance and its impact on model predictions. Two typical plots were applied: summary plots and waterfall plots. The SHAP summary plot shows the global distribution of SHAP values across all features. It presents each feature’s importance and its directional influence on predicted surgical risk. The SHAP waterfall plot breaks down individual predictions from the baseline value. It shows how each feature raises or lowers the final risk estimate. This approach supports personalized interpretation of model decisions.

Weighted Average Feature Importance Analysis and Visualization

This study combined the robustness of IG with the directional interpretability of SHAP. An equal‑weighted average of normalized importance scores was applied,31 as presented in Equation (8):

(8)

where is the combined importance score, is the min-max normalized IG score, and is the normalized absolute SHAP value (both standardized to [0,1]).

The equal-weighting scheme balances the classification discrimination ability of IG and the directional interpretability of SHAP for model behaviour. The cumulative importance percentage of the top 20 features was calculated to ensure feature dimensionality reduction and model interpretability. A weighted average importance score threshold of > 0.5 was applied to screen key predictive features, for which SHAP dependence plots were used to characterize the nonlinear relationships between each key feature and surgical intervention risk. Exploratory model-derived risk thresholds were identified using SHAP values derived from the full-feature model without retraining or feature selection, thereby providing reference for future clinical research.

Results General Clinical Data of Patients

This study collected general clinical data from 411 patients, including age, BMI, and marital status. Additionally, factors that may influence the applicability of medical abortion, such as relevant preoperative medical history, complete blood count, and biochemical markers, were also considered. Only some of these variables are presented in the following tables. Missing data were minimal and resulted from minor accidental omissions during retrospective data extraction, with exactly one missing record each for occupation, high-density lipoprotein (HDL), number of previous abortions (NPA), parity, and uterine position (UP). All missing values were imputed as detailed in the Data Preprocessing section.

In the group comparisons of Table 2 and Table 3, qualitative data were analyzed using the Chi-square test, while quantitative data were analyzed using the independent samples t-test. Among these, β-hCG, RI, RTS-FPAR, and BFS showed statistical significance (P<0.05).

Table 2 Comparison of Certain Clinical Quantitative Data Between Two Groups ()

Table 3 Comparison of Certain Clinical Qualitative Data Between Two Groups

Model Hyperparameter Settings

The final core hyperparameter configurations for each predictive model, derived from the detailed hyperparameter tuning procedure, are presented in Table 4.

Table 4 Key Parameters of the Predictive Models

Model Prediction Results

This study presents the combined results of five-fold cross-validation and hold-out validation for each model. ROC curves were plotted, and AUC (95% CI), accuracy, sensitivity and specificity were calculated to evaluate the predictive performance of different models. The detailed results of internal validation and hold-out validation are presented separately.

Internal Performance Evaluation of the Models

To visually compare the fitting performance and internal generalization ability of the models, we plotted the ROC curves for all models. Calibration plots were used to evaluate the consistency between predicted probabilities and observed event rates. Calibration plots were applied to assess how closely predicted probabilities matched observed event rates. Decision curve analysis (DCA) was used to evaluate the clinical net benefit across a range of decision thresholds. ROC, calibration, and DCA curves for the training and test sets are shown in Figure 3.

A multi plot figure with ROC, calibration and decision curve analysis graphs for training and test sets.

Figure 3 Performance evaluation of different models in internal validation. (A) ROC curve (training set). (B) ROC curve (test set). (C) Calibration curve (training set). (D) Calibration curve (test set). (E) DCA curve (training set). (F) DCA curve (test set).

In the training (Figure 3A) and test sets (Figure 3B), the stacking model showed ROC curves near the top-left corner with AUCs of 0.983 and 0.821, and outperformed LightGBM, XGBoost and other baselines, demonstrating strong discriminative ability for incomplete abortion treatment decisions. Calibration curves (Figure 3C and D) found small gaps between predicted and actual outcomes in all models, especially at low and high probabilities; the stacking model curve was close to the ideal 45° diagonal, reflecting better calibration and prediction consistency. DCA (Figure 3E and F) also verified that the stacking model yielded higher clinical net benefit than all baseline models across all threshold probabilities in both sets, supporting its better clinical utility and decision value. As evaluated by stratified 5-fold cross-validation, the detailed internal validation performance metrics of all models are presented in Table 5.

Table 5 Internal Validation Performance Metrics of All Models

Internal validation demonstrated that the ensemble models achieved promising discriminative performance for surgical intervention risk after medical abortion. The stacking model showed the best overall predictive performance, with an AUC of 0.821 (95% CI: 0.795–0.846), accuracy of 0.794, sensitivity of 0.757, and specificity of 0.806. Traditional machine learning models (LR, KNN, SVC) had relatively weaker discriminative ability, with AUC values between 0.677 and 0.735. Single ensemble models (RF, XGBoost, LightGBM, GBDT) performed moderately well, with AUC above 0.782, but were still inferior to the stacking model. These results confirm that the stacking framework can combine the strengths of different base models. It shows promise as a potential tool for assisting clinical decision-making after medical abortion.

Internal Hold-Out Performance Evaluation of the Models

Hold-out validation was performed on an independent dataset to evaluate the model’s generalization ability. Parameters optimized by internal 5-fold cross-validation were directly adopted. The discriminative performance of all models is summarized in Figure 4 and Table 6 and Table 7.

Table 6 Hold-Out Validation Performance Metrics of All Models

Table 7 Confusion Matrix Percentages (%) of All Models in Hold-Out Validation

A line graph showing receiver operating characteristic curves for multiple classification models.

Figure 4 ROC curve of different models in hold-out validation.

Hold-out validation verified that the stacking model retained the best predictive performance, with an AUC of 0.801 (95% CI: 0.700–0.892), accuracy of 0.730, sensitivity of 0.741, and specificity of 0.726. This model surpassed all other comparison models. As shown in Table 7, the confusion matrix percentages reflect balanced predictive performance for both groups: 72.60% for the non-surgical control group and 74.07% for the surgical group. These results support preliminary identification of high-risk patients who may require surgical intervention using our single-centre clinical data. Conventional machine learning models and single ensemble models showed weaker performance, with notably lower sensitivity for identifying surgical cases. These findings show moderate promising internal generalization of the stacking model within our single-centre hold-out dataset.

Feature Importance and SHAP Interpretability Analysis Results Feature Importance Analysis Based on IG and SHAP

Figure 5 presents feature importance results based on IG and SHAP values. These results were derived from weighted averaging across four ensemble models: RF, GBDT, XGBoost, and LightGBM.

Two bar graphs showing feature importance analysis based on IG and SHAP.

Figure 5 Feature importance analysis based on IG and SHAP. (A) Average IG‑based feature importance bar chart. (B) Average SHAP‑based feature importance bar chart. Bar length reflects each feature’s ability to distinguish surgical from non-surgical patients. Higher scores correspond to greater feature importance.

The results showed that β-hCG had the highest IG and acted as the core feature for model classification. Other moderately important features included BMI, age, RTS-FPAR, RI, DPLMP, and GSD-PMAUS. Their cumulative IG was above 60%, which proved their great value in differentiating the surgical group from the non-surgical group. SHAP analysis also identified β-hCG as the most influential feature, followed by RI, the first follow-up time after medical abortion (FRAM-MA), age, RTS-FPAR, GSD-PMAUS, and BMI. The core feature rankings yielded by IG and SHAP were broadly consistent in our study, reflecting the internal weight distribution of the variables in the analyzed cohort.

SHAP Visualization for Individual and Global Model Interpretation

Furthermore, we used two complementary SHAP visualization methods to explain the stacking model from individual and global perspectives, as shown in Figure 6. The SHAP waterfall plot (Figure 6A) decomposes the prediction for a representative case into contributions from individual features. Beginning at the baseline expected value , positive SHAP values (red arrows) increase the predicted risk of surgical intervention, whereas negative values (blue arrows) decrease this risk. These contributions add up to the final prediction and show the model’s decision process for this patient. The SHAP summary plot (Figure 6B) gives a global view of all study samples. Each point refers to one case: the horizontal axis is the SHAP value (positive means higher risk, negative lower risk), and color marks feature level (red for high, blue for low). This plot explains how key features affect predictions and boosts model transparency and clinical interpretability.

Side-by-side SHAP charts: a single-case contribution ladder and a feature-impact dot swarm.

Figure 6 SHAP visualization for individual and global prediction interpretation. (A) SHAP waterfall plot (individual sample interpretation). (B) SHAP summary plot (global feature interpretation).

We found RI was the most important feature driving surgical risk at the individual level. Higher RI values meant a greater chance of surgery, trailed by β-hCG, age, RTS-FPAR and other variables. For the whole study cohort, the order of feature importance differed a little. β-hCG ranked first, followed by RI, FRAM-MA, age and other factors. Combining individual-level and cohort-wide prediction patterns reveals characteristic risk associations specific to this sample and provides preliminary clues for follow-up exploratory risk assessment research.

Weighted Average Feature Importance Analysis Based on IG and SHAP

IG measures how well features distinguish different classes, while SHAP values show the size and direction of each feature’s effect. The two approaches complement each other. Figure 7 shows their weighted average to combine both approaches and evaluate feature importance more fully and reliably.

A bar chart and a line graph showing combined feature importance scores and cumulative importance percentage.

Figure 7 Weighted average feature importance analysis using IG and SHAP values. (A) Combined feature importance bar chart. (B) Cumulative feature importance curve. Higher combined scores indicate a larger contribution to predicting surgical intervention risk. The cumulative curve was used to screen key features for dimensionality reduction and improved model interpretability.

Figure 7A shows that features with combined importance scores above 0.5 were β-hCG, BMI, RI, age and RTS-FPAR. Figure 7B presents the cumulative feature importance curve. Most importance came from the top-ranked features: the top 5 contributed 58.4% of total importance, the top 8 accounted for 77.4%, and only 9 features reached the 82.0% cumulative importance threshold.

Table 8 compares the rankings of key predictive features across the three evaluation methods: IG, SHAP values, and the proposed weighted average integration. β-hCG ranked first under all three feature evaluation methods within this dataset, while BMI, RI and DPLMP exhibited varying degrees of rank discrepancy across the analytical approaches.

Table 8 Ranking Comparison of Key Predictive Features Across the Three Evaluation Methods

The weighted average ranking reduced bias from individual assessment methods. Using a threshold score above 0.5, we identified five key predictive features: β-hCG, BMI, RI, age and RTS-FPAR. These features are clinically relevant and provide quantitative support for feature selection and risk stratification in routine clinical care.

Visualization and Threshold Analysis of Key Features

We identified five key features using the IG–SHAP weighted average method. SHAP dependence plots quantified their independent effects, showed their links to surgical risk, and identified exploratory model-derived risk thresholds (Figures 8 and 9).

A scatter plot of beta hCG versus SHAP value with a rising risk trend and a marked threshold.

Figure 8 SHAP dependence plot for β‑hCG in surgical risk prediction. The x‑axis shows β‑hCG levels, and the y‑axis shows SHAP values. Positive SHAP values signify higher surgical risk. The red solid line illustrates the smoothed contribution trend of β‑hCG, and the blue dashed line marks the high‑risk threshold. Scatter point color corresponds to the predicted risk probability from the full model.

Four scatterplots showing how four patient variables shift model output, with smoothed curves and cutoffs.

Figure 9 SHAP dependence plot for predicting surgical risk. (A) BMI. (B) RTS-FPAR. (C) RI. (D) Age.

Figure 8 shows that SHAP values were mostly negative or near zero when β-hCG was below 824.90 mIU/mL, meaning little contribution to higher surgical risk. Most points were blue under multi-feature interactions, matching low surgical risk in this range. When β-hCG reached 824.90 mIU/mL or higher, SHAP values rose sharply and points turned red, indicating a substantially elevated predicted surgical risk according to our model.

Figure 9A shows that surgical risk increased significantly at BMI ≤ 18.91 or ≥23.42. Within the range of 18.91–23.42, SHAP values were close to zero and scatter points were mostly blue (low predicted risk). Fluctuations in SHAP values at low BMI may stem from the small sample size in this range or from confounding effects of mul-ti-feature interactions. At BMI ≥ 23.42, the smoothed trend line rose sharply, indicating a strong positive contribution to surgical risk that intensified with increasing BMI. However, some points remained blue, as negative effects from other key features offset the positive influence of BMI. Figure 9B demonstrated a nonlinear association between RTS-FPAR and surgical risk with marked individual variation. For RTS-FPAR ≤ 2.40 cm, it was a key determinant in patients with high SHAP values, whereas other factors (e.g, β‑hCG, RI, and clinical symptoms) predominated when SHAP values were near zero. The 2.40–3.30 cm interval had sparse data and an unstable trend, which may be limited by the study sample. Among cases with RTS-FPAR ≥ 3.30 cm, the model yielded elevated predicted surgical risk. The relatively low SHAP values in this stratum can be explained by the small number of samples with large lesion measurements. Figure 9C shows that when RI ≤ 0.51, SHAP values were mainly positive and the red trend line increased, indicating a significant positive contribution to predicted surgical risk and a greater potential need for intervention in our cohort. As RI approached 0.51, the curve declined sharply, showing that the model captured feature interactions. Even when RI had not reached its threshold, the overall risk impact of RI was altered by key factors including β‑hCG and BMI. Blue scatter points show that negative effects from other features offset the positive impact of RI. When RI was above 0.51, SHAP values were negative or close to zero, indicating a clear reduction in surgical risk, and this pattern remained relatively stable across the observed range. Under multi‑feature interactions, most points appeared blue, indicating low risk of intervention. Figure 9D shows that age-related SHAP values fluctuated slightly near zero between 20 and 40 years old. There was no clear trend toward higher or lower surgical risk. Small temporary changes included a weak negative shift at 20–22 years and positive peaks at 25–27 and 38–40 years. These patterns only caused minor fluctuations in model predictions, with no stable clinical risk signal.

Discussion

This study developed a preliminary interpretable prediction model for surgical evacuation risk in incomplete abortion based on a single-centre retrospective dataset, and identified exploratory model-derived risk thresholds to act as a candidate reference for future research. All findings are hypothesis-generating only; the potential clinical and public health benefits described below are theoretical and untested in the current study, and can only be evaluated after rigorous multi-centre external validation and decision-impact analysis. This analytical framework may also serve as a methodological reference for risk prediction research in other gynecological disorders, while cross-population generalisability remains unconfirmed.

A Preliminary Interpretable Model for Risk Stratification

Machine learning exhibits unique advantages in obstetric prediction by integrating multi-dimensional clinical data.9 This study is the first to develop a stacking ensemble model for predicting surgical intervention in incomplete abortion, whose optimized structure effectively mitigates overfitting. In internal hold-out validation, the model achieved an AUC of 0.801, with a 4.4% difference between cross-validation and hold-out validation results, indicating relatively acceptable internal consistency.32 As a preliminary exploratory candidate for risk stratification, this model may potentially reduce inconsistent clinical judgments among clinicians and across medical tiers, if validated in diverse populations. The interpretable framework based on IG and SHAP generates feature rankings specific to the present model,33,34 which offers insights into internal model behaviour and may inform subsequent external validation research.

Risk Interpretation of Major Predictors and Implications for Practice

β-hCG was ranked as the top predictor in the combined IG–SHAP importance analysis, with a risk threshold of 824.90 mIU/mL. Elevated β-hCG indicates persistent trophoblastic activity that impairs uterine contraction and the expulsion of residual tissue, corresponding to a significantly higher surgical risk.35,36 BMI also represented a key predictive factor: low BMI may reduce uterine contractility due to insufficient nutritional reserves and hormonal imbalance,37 whereas high BMI may disturb the coordination of uterine contractions and endometrial receptivity, delay tissue evacuation, and consequently increase surgical risk.38–40 As a key ultrasound marker for vascularity in intrauterine residual tissue, a lower RI suggests rich vascularization and viable residual tissue, which is associated with a greater need for surgical intervention.41 RTS-FPAR was also a statistically significant predictor of surgery (P < 0.05); larger RTS-FPAR hinders spontaneous tissue expulsion and increases the likelihood of operative management.42,43

These routinely available clinical indicators lay a practical foundation for subsequent external validation studies across different care settings. The exploratory candidate thresholds for risk stratification derived from this study may provide quantitative references for auxiliary clinical decision-making in incomplete abortion after multi-centre external validation.

Comparison with Existing Clinical Guidance and Refinement of Thresholds

Foreste et al reported that a β‑hCG level above 80 IU/L suggests increased trophoblastic activity in retained products of conception and may warrant surgical intervention

Comments (0)

No login
gif