Beyond the abdomen: an interpretable machine learning model for predicting postoperative ileus in non-abdominal surgery

Abstract

Purpose:

Postoperative ileus (POI) following non-abdominal surgery is an underestimated complication. This study aimed to develop and validate a machine learning model to predict POI risk, specifically integrating brain-gut axis variables, including depression history and chronic selective serotonin reuptake inhibitor (SSRI) use.

Methods:

A multicenter retrospective study included 2000 patients undergoing non-abdominal surgery. The cohort was divided into training (n=1050), internal testing (n=450), and external validation (n=500) cohorts. A dual-algorithm feature selection strategy combining LASSO and Boruta was used to identify robust predictors. Eight machine learning algorithms were developed and compared. Model performance was evaluated using the Area Under the Curve (AUC), calibration plots, and Decision Curve Analysis.

Results:

Seven independent predictors were identified: chronic SSRI use, history of depression, intraoperative opioids, duration of surgery, neutrophil-to-lymphocyte ratio, serum albumin, and fluid balance. The Random Forest model demonstrated superior discrimination, achieving an AUC of 0.942 in the training cohort, 0.917 in the internal testing cohort, and 0.895 in the external validation cohort. It significantly outperformed standard logistic regression (p<0.05) and displayed excellent calibration. Decision curve analysis indicated a high net clinical benefit, while SHAP analysis visually confirmed the substantial contribution of brain-gut axis factors to delayed bowel recovery.

Conclusion:

The Random Forest model provides a robust and generalizable tool for predicting POI in non-abdominal surgery patients. By highlighting the critical influence of the brain-gut axis, this study offers new insights for risk stratification and personalized perioperative management.

Introduction

Postoperative ileus is one of the most common and frequent complications following surgery (Xue et al., 2021). It is characterized by a transient impairment of bowel motility that prevents the effective transit of intestinal contents (Humphry et al., 2023). This condition leads to significant perioperative morbidity, including prolonged hospitalization, increased healthcare costs, and a higher risk of pulmonary complications and surgical site infections (Song et al., 2022). While postoperative ileus is extensively studied in the context of abdominal surgery where direct bowel manipulation occurs, its occurrence in non-abdominal surgery is often underestimated. However, the systemic stress response, anesthesia, and pharmacological interventions associated with non-abdominal procedures can still precipitate substantial gastrointestinal dysfunction, necessitating reliable risk stratification strategies for this specific patient population (Deo et al., 2022).

The pathophysiology of postoperative ileus is multifactorial, involving neurogenic, inflammatory, and pharmacological mechanisms (Jin et al., 2023). Recently, the bidirectional communication between the central nervous system and the enteric nervous system, known as the brain-gut axis, has garnered increasing attention (Meade et al., 2022). Serotonin, a key neurotransmitter regulating mood, also plays a critical role in modulating gastrointestinal motility (Chang et al., 2014). Consequently, preoperative psychological states such as depression and the chronic use of psychotropic medications like selective serotonin reuptake inhibitors (SSRI) may profoundly influence postoperative bowel recovery. Despite this biological plausibility, few predictive models have explicitly integrated these brain-gut axis variables into the risk assessment of postoperative ileus.

Traditional predictive models typically rely on conventional logistic regression analysis. While useful, these linear models often struggle to capture complex, non-linear interactions among high-dimensional clinical variables. In contrast, machine learning algorithms offer a superior alternative by efficiently processing large datasets and identifying subtle patterns that escape standard statistical methods. Furthermore, advanced feature selection techniques and model interpretation tools, such as SHapley Additive exPlanations, help bridge the gap between complex algorithmic predictions and clinical comprehensibility.

Therefore, the primary aim of this study was to develop and validate a robust machine learning model to predict the risk of postoperative ileus in patients undergoing non-abdominal surgery. We specifically sought to investigate the predictive value of brain-gut axis indicators, including a history of depression and chronic SSRI use. By employing a rigorous feature selection strategy based on the intersection of LASSO and Boruta algorithms, and by validating the model in an independent external cohort, this study intends to provide clinicians with an accurate and interpretable tool to optimize perioperative management and patient outcomes.

Materials and methodsStudy design and participants

This retrospective cohort study was conducted at two independent medical centers between January 2018 and December 2025. The study protocol was approved by the Institutional Review Board, and the requirement for informed consent was waived due to the retrospective nature of the analysis (Approval No. 2026L-03-60). We recruited patients who underwent non-abdominal surgeries under general anesthesia. Taizhou Central Hospital (Taizhou University Hospital) served as the primary center for model development, where patients were randomly divided into a training cohort and an internal testing cohort using a 70:30 ratio. The Fourth Affiliated Hospital of Soochow University served as an independent center to provide an external validation cohort, ensuring the generalizability of the developed models.

We applied strict inclusion and exclusion criteria to ensure data quality. Patients aged 18 years or older were included in the study. We excluded patients who had preoperative bowel obstruction, underwent emergency surgery, or required a second operation during the same hospitalization. Additionally, patients with significant missing data exceeding 20% of the key variables were excluded from the final analysis.

Outcome definition and data collection

The primary outcome of this study was the incidence of postoperative ileus. The diagnosis of postoperative ileus was determined through a comprehensive clinical assessment. Specifically, it was defined as the absence of flatus for more than four days after surgery, combined with clinical symptoms such as abdominal distension, nausea, and vomiting, as well as patient subjective complaints indicating bowel dysfunction (Chapman et al., 2018). This specific >4 days threshold was chosen to define clinically significant “prolonged” postoperative ileus (PPOI). According to previous global consensus and literature (Vather et al., 2013), mild bowel delays resolving within 3–4 days are often considered a transient, physiological response to anesthesia and perioperative opioids. By defining POI as >4 days, our model specifically targets patients who are at risk of prolonged recovery, which is strictly associated with increased morbidity, prolonged hospitalization, and the need for medical intervention, thereby maximizing the clinical utility of the predictive model. To minimize misclassification bias caused by varying practice patterns across centers, data extraction was highly standardized. Two independent researchers at each center systematically reviewed electronic medical records, including nursing flowsheets, physician progress notes, and medication records, using uniform diagnostic criteria. Any ambiguous clinical documentation or discrepancies between reviewers were resolved through consensus with a senior clinician.

Data were extracted from the electronic medical record systems of the participating hospitals. We collected a total of 38 potential predictor variables for each patient as detailed in Table 1. These variables included demographic characteristics, preoperative comorbidities, intraoperative parameters, and postoperative medications. Given the focus of our study on the brain-gut axis, we specifically extracted data regarding the history of depression and the chronic use of selective serotonin reuptake inhibitors to evaluate their predictive value.

VariablesRisk factorsAssignmentX1AgeContinuous variableX2GenderMale=1, Female=0X3BMIContinuous variableX4ASA ClassificationContinuous variable (1–5)X5Smoking historyYes=1, No=0X6Preoperative opioid useYes=1, No=0X7History of anxiety or depressionYes=1, No=0X8Chronic SSRI useYes=1, No=0X9Preoperative sleep quality score (PSQI)Continuous variableX10Preoperative cognitive score (MMSE)Continuous variableX11History of motion sicknessYes=1, No=0X12Intraoperative BIS value (Mean)Continuous variableX13Diabetes MellitusYes=1, No=0X14Hypertension or CVDYes=1, No=0X15COPDYes=1, No=0X16History of neurological diseaseYes=1, No=0X17Serum AlbuminContinuous variableX18Neutrophil-to-lymphocyte ratio (NLR)Continuous variableX19C-reactive protein (CRP)Continuous variableX20HemoglobinContinuous variableX21Preoperative potassiumContinuous variableX22eGFRContinuous variableX23Fasting blood glucoseContinuous variableX24Duration of surgeryContinuous variableX25Duration of anesthesiaContinuous variableX26Spine surgeryYes=1, No=0X27Intraoperative opioids (MME)Continuous variableX28Anesthesia maintenance typeInhalational=1, TIVA = 0X29Duration of hypotension (MAP < 65mmHg)Continuous variableX30Vasopressor useYes=1, No=0X31NMBA reversal agentSugammadex=1, Neostigmine=0X32Fluid balanceContinuous variableX33Estimated blood lossContinuous variableX34Postoperative PCIA useYes=1, No=0X35Pain score at POD1 (VAS)Continuous variableX36Time to first mobilizationContinuous variableX37Postoperative potassiumContinuous variableX38Perioperative blood transfusionYes=1, No=0YPostoperative IleusYes=1, No=0

BMI, body mass index; ASA, American Society of Anesthesiologists; SSRI, selective serotonin reuptake inhibitors; PSQI, Pittsburgh sleep quality index; MMSE, mini-mental state examination; BIS, bispectral index; CVD, cardiovascular disease; COPD, chronic obstructive pulmonary disease; NLR, neutrophil-to-lymphocyte ratio; CRP, C-reactive protein; eGFR, estimated glomerular filtration rate; MME, morphine milligram equivalents; TIVA, total intravenous anesthesia; MAP, mean arterial pressure; NMBA, neuromuscular blocking agent; PCIA, patient-controlled intravenous analgesia; POD1, postoperative day 1; VAS, visual analog scale.

Data preprocessing and feature selection

Prior to analysis, the raw data underwent preprocessing to handle missing values and standardize the format. After excluding patients with significant missingness (>20%), the missing rates for all remaining individual variables were extremely low (all < 5%). The detailed missing rates for each variable are reported in Supplementary Table 2. An analysis of the missing data patterns suggested that the mechanism of missingness was predominantly Missing Completely at Random (MCAR) or Missing at Random (MAR), primarily attributable to random administrative omissions in the electronic medical records. For continuous variables with missing data, we employed mean imputation, while missing categorical values were filled using the mode imputation method. This single-imputation approach was chosen for its computational efficiency and to streamline the preprocessing pipeline across multiple machine learning algorithms. Its use was justified by the relatively low overall proportion of missing data, as patients with extensive missingness (>20%) had already been excluded. Continuous features were subsequently standardized using Z-score normalization to ensure that all variables contributed equally to the model.

We employed a rigorous two-step feature selection strategy to identify the most relevant predictors and eliminate redundancy. First, we utilized the Least Absolute Shrinkage and Selection Operator regression algorithm, which applies a penalty to shrink coefficients and select variables. Second, we applied the Boruta algorithm, a random forest-based method that compares original features against randomized shadow features to determine statistical significance. The final set of predictors was determined by taking the intersection of the variables selected by both the LASSO and Boruta algorithms. To ensure the independence of the selected features, we performed both Spearman correlation analysis and Variance Inflation Factor assessment. This dual approach allowed us to rigorously rule out multicollinearity among the predictors.

Machine learning model development

We developed and compared eight distinct machine learning algorithms to predict the risk of postoperative ileus. These algorithms included Random Forest, XGBoost, LightGBM, Logistic Regression, Naive Bayes, Decision Tree, Support Vector Machine, and K-Nearest Neighbors. To optimize the performance of each model, we conducted hyperparameter tuning using grid search combined with 10-fold cross-validation on the training dataset. This process involved systematically testing various combinations of parameters to identify the configuration that yielded the best predictive accuracy while preventing overfitting.

Model evaluation

The performance of the models was evaluated comprehensively in the training, internal testing, and external validation cohorts. We primarily used the Receiver Operating Characteristic curve and the Area Under the Curve to assess the discrimination ability of the models. We also calculated other key performance metrics including sensitivity, specificity, accuracy, and the F1-score. In addition to discrimination metrics and visual calibration curves, the Brier score was calculated for all evaluated models as a quantitative measure of calibration. The Brier score evaluates the mean squared difference between the predicted probabilities and the actual observed outcomes, with a lower score (closer to 0) indicating superior calibration.

Statistical comparisons between the Area Under the Curve values of different models were performed using the DeLong test, with a two-sided P value of less than 0.05 considered statistically significant. Furthermore, we assessed the calibration of the models using calibration curves to examine the agreement between predicted probabilities and observed outcomes. The clinical utility of the models was evaluated using Decision Curve Analysis, which quantifies the net benefit of using the model across a range of threshold probabilities.

Model interpretability

To address the “black box” nature of machine learning models and enhance clinical transparency, we utilized SHapley Additive exPlanations. This method assigns an importance value to each feature for a particular prediction. We generated global importance plots to rank the overall contribution of each variable and summary plots to visualize the direction and magnitude of the feature effects. Additionally, waterfall plots were created to demonstrate how specific risk factors contributed to the predicted probability for individual patients, facilitating personalized clinical interpretation.

Statistical analyses

All statistical analyses and data visualizations were performed using.

R (version 4.4.2). Continuous variables were assessed for normality via the Shapiro-Wilk test. Normally distributed data are presented as mean ± standard deviation, with group comparisons conducted using Student’s t-tests. Non-normally distributed variables are expressed as median and interquartile range [M (Q1, Q3)] and analyzed via the Mann-Whitney U test. Categorical variables are reported as frequencies (percentages) and evaluated using Chi-square tests or Fisher’s exact tests (for cell counts <5). Statistical significance was defined as a two-tailed p-value < 0.05.

ResultBaseline characteristics and variable definitions

A total of 2000 patients undergoing non-abdominal surgery were included in this study and randomly allocated into three distinct cohorts. These consisted of a training set with 1050 patients, an internal testing set with 450 patients, and an external validation set with 500 patients. Table 2 details the baseline demographic and clinical characteristics. The average age of patients in the training cohort was 59.4 ± 11.5 years, which was comparable to the testing and validation cohorts with P values exceeding 0.05. The overall incidence of postoperative ileus was 15.4% in the training set, 16.9% in the testing set, and 14.2% in the validation set, showing no statistically significant difference with a P value of 0.485. Similarly, key brain-gut axis variables such as the history of depression and chronic SSRI use were balanced across groups, with prevalence rates of approximately 13.5% and 6.5% in the training cohort respectively. These results confirm that the randomization process effectively balanced the baseline features.

CharacteristicTraining cohort (n = 1050)Testing cohort (n = 450)Validation cohort (n = 500)P valueDemographicsAge (years), Mean ± SD59.4 ± 11.560.8 ± 10.958.7 ± 12.10.076Gender (Male), n (%)512 (48.8%)234 (52.0%)229 (45.8%)0.143BMI (kg/m²), Mean ± SD24.5 ± 3.224.9 ± 3.524.2 ± 3.10.089ASA classification, n (%)0.215I - II785 (74.8%)318 (70.7%)382 (76.4%)III - IV265 (25.2%)132 (29.3%)118 (23.6%)Smoking history (Yes), n (%)258 (24.6%)129 (28.7%)110 (22.0%)0.094Preoperative opioid use (Yes), n (%)86 (8.2%)45 (10.0%)34 (6.8%)0.187Gut-brain axis factorsHistory of depression (Yes), n (%)142 (13.5%)74 (16.4%)61 (12.2%)0.156Chronic SSRI use (Yes), n (%)68 (6.5%)36 (8.0%)28 (5.6%)0.342Sleep quality score (PSQI), Mean ± SD7.4 ± 3.57.8 ± 3.87.2 ± 3.30.128Cognitive score (MMSE), Mean ± SD27.5 ± 2.127.2 ± 2.427.6 ± 1.90.203History of motion sickness (Yes), n (%)195 (18.6%)75 (16.7%)108 (21.6%)0.165Intraoperative BIS (Mean), Mean ± SD48.5 ± 5.647.9 ± 6.149.1 ± 5.20.068ComorbiditiesDiabetes Mellitus (Yes), n (%)176 (16.8%)88 (19.6%)75 (15.0%)0.189Hypertension/CVD (Yes), n (%)438 (41.7%)199 (44.2%)195 (39.0%)0.274COPD (Yes), n (%)65 (6.2%)36 (8.0%)26 (5.2%)0.211Neurological disease history (Yes), n (%)48 (4.6%)28 (6.2%)19 (3.8%)0.235Preoperative labsSerum Albumin (g/L), Mean ± SD39.5 ± 4.239.1 ± 4.539.8 ± 4.00.154NLR, Mean ± SD2.45 ± 1.122.58 ± 1.252.39 ± 1.050.092CRP (mg/L), Median (IQR)3.5 (1.8-6.2)3.9 (2.0-6.8)3.2 (1.6-5.9)0.118Hemoglobin (g/L), Mean ± SD132.5 ± 15.6129.8 ± 16.4133.4 ± 14.80.072Preoperative potassium (mmol/L), Mean ± SD4.05 ± 0.354.02 ± 0.384.08 ± 0.320.165eGFR (mL/min), Mean ± SD88.5 ± 18.286.4 ± 19.589.2 ± 17.60.246Fasting blood glucose (mmol/L), Mean ± SD5.8 ± 1.56.1 ± 1.85.7 ± 1.40.082Intraoperative variablesDuration of surgery (min), Mean ± SD165.4 ± 45.2172.1 ± 48.6160.5 ± 42.80.058Duration of anesthesia (min), Mean ± SD205.4 ± 55.2214.5 ± 58.6199.2 ± 50.80.064Spine surgery, (Yes),n (%)388 (37.0%)185 (41.1%)172 (34.4%)0.096Intraoperative opioids (MME, mg), Mean ± SD45.8 ± 12.547.2 ± 13.844.9 ± 11.60.137Anesthesia maintenance, n (%)0.312Inhalational651 (62.0%)268 (59.6%)322 (64.4%)TIVA399 (38.0%)182 (40.4%)178 (35.6%)Duration of hypotension (min), Median (IQR)10 (0-25)12 (0-28)8 (0-22)0.145Vasopressor use (Yes), n (%)345 (32.9%)162 (36.0%)151 (30.2%)0.178NMBA Reversal (Sugammadex), n (%)498 (47.4%)230 (51.1%)228 (45.6%)0.254Fluid balance (mL), Mean ± SD1250 ± 4501310 ± 4801215 ± 4200.063Estimated blood loss (mL), Median (IQR)200 (100-350)220 (120-380)190 (100-320)0.088Postoperative outcomesPCIA use (Yes), n (%)815 (77.6%)338 (75.1%)402 (80.4%)0.162Pain score at POD1 (VAS), Mean ± SD3.4 ± 1.23.6 ± 1.43.3 ± 1.10.065Time to first mobilization (h), Mean ± SD22.5 ± 6.823.4 ± 7.521.9 ± 6.20.074Postoperative potassium (mmol/L), Mean ± SD3.85 ± 0.423.81 ± 0.453.88 ± 0.400.193Perioperative transfusion (Yes), n (%)115 (11.0%)62 (13.8%)48 (9.6%)0.104Outcome: Postoperative Ileus, (Yes), n (%)162 (15.4%)76 (16.9%)71 (14.2%)0.485

Baseline characteristics of patients in the training, testing, and validation cohorts.

Data are presented as n (%) for categorical variables and as mean ± SD or median (IQR) for continuous variables, as appropriate. The symbol (%) denotes the percentage of patients within the corresponding cohort.

BMI, body mass index; ASA, American Society of Anesthesiologists; SSRI, selective serotonin reuptake inhibitors; PSQI, Pittsburgh sleep quality index; MMSE, mini-mental state examination; BIS, bispectral index; CVD, cardiovascular disease; COPD, chronic obstructive pulmonary disease; NLR, neutrophil-to-lymphocyte ratio; CRP, C-reactive protein; eGFR, estimated glomerular filtration rate; MME, morphine milligram equivalents; TIVA, total intravenous anesthesia; MAP, mean arterial pressure; NMBA, neuromuscular blocking agent; PCIA, patient-controlled intravenous analgesia; POD1, postoperative day 1; VAS, visual analog scale.

We collected 38 potential predictors for each patient to construct the model. Table 1 provides a comprehensive list of these variables along with their assignment criteria. The candidate features encompassed diverse categories including demographic factors like age and gender, preoperative comorbidities such as diabetes and hypertension, and specific intraoperative metrics including the duration of surgery and intraoperative opioid dosage. Categorical variables such as gender and smoking history were encoded as binary values, while continuous variables like BMI and laboratory indicators retained their original numerical measurements.

Feature selection via intersection of algorithms and collinearity check

To isolate the most robust predictors from the 38 candidate variables, we implemented a dual-algorithm selection strategy. The Least Absolute Shrinkage and Selection Operator regression was first applied to minimize overfitting. As illustrated in Figure 1A, the coefficient profiles of the variables converged as the penalty parameter increased. The optimal lambda value was determined via 10-fold cross-validation as shown in Figure 1B. Concurrently, the Boruta algorithm was utilized to evaluate feature importance relative to randomized shadow attributes. Figure 1C depicts the selection process over multiple iterations, and Figure 1D highlights the final confirmed attributes that performed significantly better than the shadow maxima.

Panel A shows a line plot illustrating coefficient trajectories against log lambda values for feature selection in a penalized regression model. Panel B presents a line graph of binomial deviance versus log lambda with minima and error bars for model validation. Panel C depicts a line plot of variable importance scores across classifier runs, where multiple features are tracked in green, red, and blue. Panel D is a horizontal bar chart ranking attribute importance for various clinical variables, with colored bars indicating importance distribution and confidence intervals. Panel E displays a correlation matrix heatmap of seven selected clinical features, with color intensity indicating correlation strength and direction.

Feature selection and correlation analysis. (A) LASSO coefficient profiles of candidate variables. (B) Selection of the optimal lambda (λλ) in the LASSO model using 10-fold cross-validation. (C) Dynamic feature importance evolution during the Boruta algorithm process. (D) Final feature importance ranking by Boruta; green box plots indicate confirmed predictors. (E) Spearman correlation heat map of the seven selected variables.

We determined the final model predictors by identifying the intersection of variables selected by both LASSO and Boruta. This rigorous process yielded seven core variables: history of depression, chronic SSRI use, serum albumin, neutrophil-to-lymphocyte ratio, duration of surgery, intraoperative opioids, and fluid balance. Following selection, we assessed multicollinearity to ensure model stability. Figure 1E displays the Spearman correlation heat map, which indicates weak correlations among the selected features. This was further quantified in Table 3, where the Variance Inflation Factor analysis showed that all seven variables had VIF values well below the threshold of 5. For instance, chronic SSRI use had a VIF of 1.70 and the duration of surgery had a VIF of 1.62, confirming their independence and suitability for multivariate modeling.

VariableVIF (All candidate variables)VIF (Selected features)Inclusion in final modelAge1.4528394726101Gender1.1509384756201BMI1.3247502938471ASA classification2.1503928475629Smoking history1.5203948572610Preoperative opioid use1.6829301928374History of depression1.20583746291031.0982374610293YesChronic SSRI use1.48291029384721.7046186409865

Comments (0)

No login
gif