This multicenter observational, retrospective, cohort study identified patients with shock between January 1st, 2018, and December 31st, 2023, within an integrated academic urban healthcare system located in the Midwestern United States comprised of 9 acute care hospitals. These hospitals included one academic medical center, seven community hospitals, and one critical access hospital. This study was approved by the Northwestern University Institutional Review Board under study number STU00218401. The transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD) reporting guidelines was followed. All procedures were followed in accordance with the Helsinki Declaration of 1975.
Data sourceElectronic Health Record (EHR) structured and unstructured data were obtained from the healthcare system Enterprise Data Warehouse (EDW). The systemwide integrated EDW captured all inpatient encounters across all hospitals in the healthcare system from 2018 until 2023. This data source has been used to successfully identify the presence of shock as well as the shock phenotype among injured patients in a variety of geographic context (urban, suburban and rural) and hospital settings (academic medical center, community hospitals and critical access hospitals) [10].
Study populationEligibility criteria were adult patients greater than or equal to 18 years of age, hospitalized in an intensive care unit (ICU) at any one of the 9 acute care hospitals between 2018 and 2023 with shock. Shock was defined using both structured (vitals and labs) and unstructured (notes) data. Shock patient encounters were first identified with structured data, specifically a systolic blood pressure measurement of less than 90 mmHg, AND with at least one new organ failure as measured by a new elevation in the Sequential Organ Failure Assessment (SOFA) score from the last recorded measurement [11, 12]. Among these patient encounters, the shock cohort was further narrowed to include only patient encounters with “shock” documented in history and physical, progress, or consult notes using regular expression from the same day (day 0) as systolic blood pressure measurement of less than 90 mmHg and new elevation in the SOFA score. Patient encounters with notes only using negation terms were excluded. Patient encounters with less than two clinician notes documenting “shock” were excluded (eFig. 1) as two or more notes on any one day were needed to determine inter-clinician diagnostic agreement.
Shock etiology diagnosisShock etiology diagnosis was a priori collapsed into nine different shock etiologies Septic, Cardiogenic, Hypovolemic, Adrenal, Neurogenic, Undifferentiated, Obstructive, Anaphylactic, Post-procedural) based on clinical experience and prior literature.[13] Regular expressions were applied to all unstructured data documented in clinician history and physical, progress, or consult notes from the first day of the systolic blood pressure measurement of less than 90 mmHg and new elevation in the SOFA score (called day 0) through day four. A daily clinician-specific shock etiology diagnosis variable was created for each clinician note [14]. Septic shock was identified by presence of the string value “Sepsis”, “Septic”, and “Septic Shock” without negation statements. Cardiogenic shock was identified by presence of the string value “Cardiac Shock”, “Cardiogenic”, “Cardiogenic Shock” without negation statement. Hypovolemic shock was identified by presence of the string value “hypovolemic”, “hypovolemia”, “hypovolemic shock”, “hemorrhage”, “hemorrhagic shock” and “traumatic shock” without negation statement. Adrenal shock was identified by presence of the string value “adrenal shock”, and “adrenal insufficiency” without negation statement. Neurogenic shock was identified by presence of the string value “spinal shock”, and “neurogenic shock” without negation statement. Undifferentiated shock was identified by the presence of the string value “undifferentiated shock”, “shock NOS”, and “multifactorial shock” without negation statement. Obstructive shock was identified by presence of the string value “obstructive shock”, “cardiac tamponade”, “tamponade”, “tension pneumothorax” and “tension pneumo” without negation statement. Anaphylactic shock was identified by the presence of the string value “anaphylaxis”, “anaphylactic” and “anaphylactic shock” without negation statement. Post-procedural shock was identified by presence of the string value “post-procedure shock”, and “post-procedural shock” without negation statement. Shock etiology diagnosis was abstracted from each clinician note in the electronic health record, for each patient, on each day 0 through day 4. We manually reviewed all occurrences to identify if ambiguous or overlapping terms were present. “Undifferentiated Shock” was the ambiguous term most frequently used, and we created a separate vector (of the nine) labeled as such. Overlapping terms, such as the different shock diagnoses, in the same note or by the same clinician in a later note each day were treated as discrete observations generating an additional vector for each separate shock diagnosis. Additionally, we manually reviewed all observations and ensured that context-sensitive language was not playing any role here. A total of 395 patients were reviewed with ambiguous phrases adjudicated by two clinicians (LJ, AMS). Ambiguous or mixed phrases (e.g., “distributive shock,” “shock of unclear etiology”) were mapped onto the unknown category. When two different forms of shock were hypothesized then (e.g., “mixed septic and cardiogenic shock,”), the clinician note was annotated for both Septic and Cardiogenic shock.
Outcome of interestThe primary outcome of interest was inter-clinician diagnostic agreement of shock etiology diagnosis. Each clinician note, each dayfor each patient was interpreted as a nine-dimensional vector of binary values (i.e., 0 = not mentioned, 1 = mentioned) indicating mention of each of the nine shock etiologies above. Each clinicians’ daily diagnosis documented in their notes were compared to every other clinician’s daily diagnosis from days 0 to 4. The agreement between the clinicians’ shock etiology diagnoses was quantified using cosine similarity scores for nine-dimensional vectors from days 0 to 4. Cosine similarly measures the alignment of two vectors, with scores ranging from 0 to 1; a score of 1 indicates complete alignment, representing complete inter-clinician diagnostic agreement [15]. Patients were categorized based on whether they ever achieved a daily average cosine similarity score of 1 during the first four days of shock. Two near-perfect cosign similarity scores thresholds (≥ 0.95; ≥ 0.90) were also tested in sensitivity analyses. The measure of agreement was based on cosine similarity of etiology vectors, not on direct clinical adjudication. This study sought to understand the patient and clinician characteristics associated with not having complete inter-clinician diagnostic agreement of shock etiology, and to evaluate the ability of machine learning methods to predict which patients would never achieve inter-clinician agreement.
Independent variablesIndependent variables were patient demographics, clinical factors, clinical course, and clinician characteristics. Patient demographics included age in years which was categorized into five groups, 50 and under, 51–60, 61–70, 71–80 and 81 and older. Sex was categorized as male and female. Race was categorized into seven groups as White, Black/African American, Asian, Native Hawaiian/Pacific Islander, American Indian/Alaska Native, Other and Unknown. Ethnicity was categorized into three groups as Hispanic, not Hispanic, and Unknown. Insurance was categorized into four groups as Medicare, Medicaid, Private and Self-pay. Area deprivation index (ADI) rank was calculated based on the patients address [16]. A larger number suggested patient lived in a neighborhood with greater socioeconomic deprivation. The ADI rank at the state was categorized into four groups as 1–2, 3–5, 6–10 and unknown, as well as categorized by the national rank into five groups as < 25, 25–49, 50–74, 75–100 and unknown.
Clinical features were SOFA score, Sepsis Score, Shock Index, shock etiology, Elixhauser comorbidity index, and Body Mass Index. Contextual clinical features were whether the first shock diagnosis occurred during the weekday or weekend, whether the ICU admission was on the weekday or weekend, and number of unique clinicians who wrote a note on the patient in the first 4 days of shock. The first recorded SOFA score was calculated from laboratory and clinical data for each patient upon shock presentation and was categorized into five groups as 0–6, 7–12, 13–15, 16–24, no SOFA score [17]. The sepsis score was categorized into six groups as 0, 1, 2, 3, > = 4, not applicable [18]. The shock index was calculated by dividing heart rate by systolic blood pressure then categorized in five groups as < 0.5, 0.5–0.7, 0.7–1.0, > 1.0, unknown [19]. The nine shock etiologies were described above. Elixhauser comorbidity index was calculated from the International Classification of Disease Codes and categorized into four groups: 0–5, 5–6, 7–10 and > 10. Body mass index (BMI) was calculated as weight (kg) divided by height (m2) and categorized into six groups: < 18.5, 18.5–24.9, 25.0–29.9, 30.0–34.9, ≥ 35, and unknown. ICU admission day and first diagnosis of shock day were each categorized as weekday or weekend. The daily note count, and unique clinician count per patient were calculated and treated as continuous variables.
Clinical course characteristics were SOFA score changes, treatments administered, ICU mortality and discharge disposition. SOFA score change was calculated as the difference between day 0 and day 4 then was categorized into four groups: improved, no change/worse, no SOFA score, and only one value recorded. Treatments included time to first intravenous bolus fluids (normal saline, lactated ringers or plasmalyte) and was categorized as < 24, 24–48, 49–72, 73–96, and > 96 h. The administration of intravenous fluids was recorded dichotomously by the day of shock presentation (days 0–4). Antifungal administration (micafungin, fluconazole, amphotericin) was categorzied as a dichotomous variable. Time to first antibiotics (eTable 1) was categorized similarly to intravenous fluids. Anticoagulation infusions (heparin, bivalirudin, argatroban) and diuretics (furosemide, bumetanide) were categorized dichotomously. Time to first lab draw, echocardiogram, and surgical procedure were categorized into six group variable: < 24, 24–48, 49–72, 73–96, and > 96 h. Procedures such as arterial line placement, cardiac catheterization, central venous catheter placement, and continuous renal replacement therapy were recorded as dichotomous variable . ICU mortality was a dichotomous outcome Discharge disposition was categorized as transfer to another acute care hospital, acute inpatient rehabilitation, long-term care facility, home, death, or unknown. Clinician characteristics included clinician type (e.g., physician, advanced practice provider) and specialty (e.g., cardiology).
Statistical analysisThe rate of patients which never had complete inter-clinician diagnostic agreement on any of the four days was calculated. Differences in patient demographics, clinical features, and clinical course between patients with complete inter-clinician diagnostic agreement and without were compared using chi-squared tests of independence. Clinician characteristics were compared between notes with and without complete diagnostic agreement using chi-squared tests, with statistical significance set at p-value < 0.05. Unadjusted analyses were performed in R version 4.3.0.
Machine learning models were developed to predict patients without complete inter-clinician diagnostic agreement on shock etiology. Data were split into 80% training and 20% validation cohorts. On the 80% training data, we performed fivefold cross-validation to tune the parameters of each of the six machine learning models. The six machine learning models tuned were logistic regression, Random Forest, Decision Trees, Gradient Boosting, K-Nearest Neighbor (KNN), and Support Vector Machine (SVM) algorithms. The best-performing parameters from cross-validation were used to train on all 80% of data to obtain the final models. These models aimed to predict patients without complete inter-clinician diagnostic agreement. Model performance was assessed using accuracy, F1 score, and the receiver operating characteristic (AUC-ROC) on the validation cohort, with metrics ranging from 0.0 to 1.0, where values closer to 1.0 indicate higher performance [20]. Calibration plots were created to compare predicted and observed probabilities for each of the six models. A perfectly calibrated model’s curve should align closely with the diagonal, indicating accurate predictions. The hyperparameters that were tuned were the AUC-ROC and F1 score using the fivefold cross-validation on the 80% training set only. The reported performance metrics (accuracy, F1, AUC-ROC) and calibration curves were all based on the 20% held-out validation cohort. Two sensitivity analyses were conducted to evaluate overly stringent threshold for inter-clinician agreement and model evaluation was performed for near-perfect cosign similarity scores thresholds (≥ 0.95; ≥ 0.90).
Feature importance was calculated using Random Forest algorithm to identify patient, clinical, and clinician features predictive of inter-clinician diagnostic agreement on shock etiology. Random Forest algorithm ranks feature importance in decision tree regression [21]. Analyses were performed in Python version 3.10.9.
Comments (0)