AI and Internet of Things for Chronic Obstructive Pulmonary Disease Remote Monitoring: Systematic Review of Exacerbation Prediction and Key Physiological Variables


Introduction

Chronic obstructive pulmonary disease (COPD) is a heterogeneous lung condition characterized by chronic respiratory problems, including dyspnea, cough, sputum production, and exacerbations. These issues can be attributed to abnormalities of the airways (bronchitis and bronchiolitis) and/or alveoli (emphysema) leading to persistent, often progressive, airflow obstruction []. COPD is a major global health issue, causing significant morbidity and mortality worldwide, and placing an increasing burden on health care systems and the economy. The current global estimated prevalence of 10.3% is projected to rise, particularly in high-income countries, mainly due to an aging population []. According to the World Health Organization, COPD was responsible for 3.5 million deaths in 2021 alone, making it the fourth leading cause of death worldwide [].

The primary factors contributing to disease progression, hospitalization, and readmission are exacerbations of COPD (ECOPD), which refers to acute episodes of disease worsening. According to the Global Initiative for Chronic Obstructive Lung Disease (GOLD) definition, these are defined as “an event characterized by increased dyspnea and/or cough and sputum that worsens in less than 14 days, which may be accompanied by tachypnea and/or tachycardia and is often associated with increased local and systemic inflammation caused by infection, pollution, or other insults to the airways” [,]. ECOPD episodes are classified as mild or moderate, according to the medication treatment, or severe, if hospitalization or emergency room (ER) visits are required []. Thus, the severity of ECOPD is graded post hoc, based on the medication used and the setting, and is determined by the patient’s subjective perception of symptoms.

Currently, there is no single universally recognized gold standard to assess ECOPD. Instead, exacerbations are commonly identified through a combination of symptom-based or event-based assessments, patient self-reports, and clinical evaluation []. ECOPD is defined as event-based, using objective criteria such as health care use or specific treatments, including self-administration of medication or unscheduled visits to the ER and/or hospital admission. Alternatively, it can be defined as symptom-based, relying on patient-reported changes in symptoms. An increase in symptoms such as dyspnea, cough, or sputum production for a minimum number of consecutive days, or changes in symptoms recorded in daily diaries or questionnaires. However, it is logical that the symptoms precede the event, as they are the underlying cause.

It is estimated that approximately 30% to 50% of people diagnosed with COPD experience at least one exacerbation yearly, often requiring hospitalization to stabilize the patient []. A significant concern following these hospitalizations is the high rate of readmissions. Data from the European COPD audit revealed that 35.1% (5337/15,191) of patients were readmitted within 90 days of discharge []. ER visits and hospitalizations due to signs and symptoms of ECOPD are the main reason for COPD-related costs, which can account for up to 56% of the annual respiratory disease budget, equating to €38.6 billion (US $45.7 billion) [,]. A qualitative study by Locke et al [] found that most patients can distinguish worsening disease symptoms from baseline symptoms, but a reluctance to seek care was a typical response of many persons with COPD when developing ECOPD. Most delayed action until a critical threshold was reached, when the urgency of their symptoms surpassed the personal and initial barriers they faced. In contrast, early recognition and prompt treatment by a physician can improve the outcomes of ECOPD []. These findings underscore the importance of implementing strategies aimed at early recognition of disease worsening, enabling prompt therapeutic action.

To achieve effective early prediction, it is essential to identify key parameters for monitoring and establish consensus on alerting thresholds. This process involves identifying patients at high risk before the onset of severe symptoms. Machine learning (ML) models, particularly those integrated with emerging Internet of Things (IoT) technologies, offer promising solutions for accurately and remotely sensing vital parameters. Equally important is determining the optimal time window for early prediction. This window should allow for sufficient lead time to initiate interventions that could prevent or attenuate an exacerbation. If the time frame is too short, there may not be sufficient time to act effectively. Conversely, a time frame that is too long, such as forecasting ECOPD risk over the next 12 months, may have limited clinical utility, as it does not inform the timing of specific interventions. Therefore, determining the optimal lead time for prediction is critical for maximizing both predictive accuracy and clinical actionability. This choice must balance the need for early warning with the relevance and reliability of available data in the lead-up to an exacerbation.

Beyond achieving accurate early prediction, ensuring clinical translation and adoption is equally important. Recent reviews, such as Qi et al [], underscore the importance of structured validation frameworks and clinical interpretability as essential prerequisites for integration into clinical practice. Consequently, these considerations are as important as the previously discussed factors in the development of early prediction models. In particular, it is essential to ensure that the model is both generalizable and explainable for successful clinical adoption.

From a practical perspective, ECOPD prediction involves several fundamental steps, including the identification of physiological variables and external factors that change and/or contribute to the lead-up to an exacerbation. It also involves the selection of wearable and/or portable devices to collect these variables. It is then important to identify key parameters that change during prodromal phases, that is, before exacerbation of symptoms, as such information may positively influence model performance. The collected data are then preprocessed using various techniques and subsequently fed into artificial intelligence (AI)/ML algorithms. The objective of the algorithm could be early warning generation (with varying prediction windows), the detection of exacerbation onset, or the classification of prodromal phases. The entire ECOPD prediction pipeline is illustrated in .

Figure 1. Schematic overview of exacerbations of chronic obstructive pulmonary disease prediction pipeline. AI: artificial intelligence; COPD: chronic obstructive pulmonary disease; ECOPD: exacerbations of chronic obstructive pulmonary disease.

The goal of this systematic review is to analyze existing research on AI- and IoT-based wearable systems for remote COPD monitoring, focusing on which variables are evaluated to assess and/or predict exacerbations, the granularity of the collected data, and the device type used to enable remote patient monitoring (RPM). It highlights state-of-the-art results on prediction accuracy and area under the curve (AUC), outlining which AI and ML techniques are most commonly used. It focuses on examining existing data frameworks, the AI models used for predicting ECOPD, and identifying challenges in their real-world application, as well as gaps in the current research.

The research questions can be summarized as follows:

Which key physiological variables are most relevant for predicting COPD exacerbations? Does the data resolution (eg, sampling frequency) influence model performance? Is there any standardization in the IoT devices used for RPM in COPD management?Which AI models are used to predict ECOPD, and how accurate are they?What are the gaps in current research regarding wearable IoT and AI for COPD monitoring, and what should be prioritized in future studies?

The first research question aims to identify which physiological and environmental variables are most commonly used, exploring the frequency of data collection and its impact on the outcome. It further reviews the types of wearables and IoT devices used, and whether there is any standardization across studies. The second research question discusses the ML techniques applied, including performance metrics and limitations. It also clarifies how exacerbations are defined, which influences model training and evaluation. Finally, the third one synthesizes insights from all thematic areas to identify limitations in current approaches and propose future directions. Each of these points is examined thoroughly in the “Discussion” section.

To support this analysis, we also categorize the reviewed articles based on their primary objectives. Studies that focused on developing models for predicting exacerbations are discussed in the section titled “Exacerbation Prediction,” while those that examined changes in vital parameters preceding exacerbations are presented in the section titled “Variable Changes in Prodromal Phases.” This second group of papers was deliberately included to highlight the importance of identifying discriminative input variables, which is critical for improving model performance and capturing the complex patterns of disease progression. Despite this thematic separation, the discussion of results remains structured around the 3 central research questions, ensuring a coherent and focused analysis.

Among the literature on reviewing ECOPD, this review adds key insights for developing an effective remote system for early ECOPD prediction. We identify monitored variables that have discriminative power to predict ECOPD, underscore the need for a universal definition of ECOPD, and review state-of-the-art telemonitoring solutions and early prediction methods. By doing so, this review aims to lay a solid foundation for improving outcomes and their validation within the context of RPM for ECOPD. This review aims to emphasize the necessity of future research in 3 key areas: reducing the burden on medical specialists through automated data detection and prediction, alleviating patient burden by minimizing active participation in data collection, and enhancing predictive models through multimodal strategies.


MethodsEligibility Criteria, Information Sources, and Search Strategy

The review was conducted in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines []. A comprehensive search was carried out in PubMed, Scopus, IEEE Xplore, and Embase during March 2025. The search strings can be found in .

Only peer-reviewed, full-text articles published in English were considered. The following publication types were excluded: systematic reviews, meta-analyses, book chapters, editorials, letters, and conference abstracts. However, full papers in conference proceedings were included.

Studies were eligible for inclusion if they met the following criteria:

Population: patients with COPD, undergoing exacerbations of the condition.Settings: RPMMeasurements: objective (ie, measurable) physiological variables

The search strategy was based on three core concepts:

Wearable and remote monitoring (RM) technologies (eg, wearable sensors, smartwatches, home monitoring, and IoT sensors)COPDML and AI approaches for the prediction or detection of ECOPD

Both controlled vocabulary terms (eg, MeSH [Medical Subject Headings] in PubMed) and free-text keywords were used. The search strategies were adapted for each database’s specific syntax and indexing system.

No additional filters or limitations were applied during the search process.

Selection Process and Data Collection

The screening process was conducted using Rayyan (Qatar Computing Research Institute) [], a web-based tool for systematic review management. Duplicate records were first identified automatically by the platform, but their final removal was performed manually by the first reviewer to ensure accuracy.

Two reviewers (MM and JG) independently screened the articles, initially based on titles and abstracts, followed by full-text screening for eligibility. In cases of disagreement, conflicts were resolved through discussion until consensus was reached. Articles were selected for inclusion if they met all predefined eligibility criteria.

Following the screening, the first reviewer (MM) manually extracted the relevant data using Microsoft Excel. The second reviewer (JG) subsequently verified the data for completeness and accuracy.

The following data items were extracted: title, authors, year of publication, journal, aim of the study, population characteristics, study duration, definition of exacerbation, monitored parameters, devices used, sampling frequency, machine learning or deep learning models used, and main outcomes.

Risk of Bias Assessment

The risk of bias was assessed by one reviewer (MM) and verified by a second (JG). Any disagreements between reviewers were resolved through discussion to reach consensus. To ensure a comprehensive evaluation, several risk of bias assessment tools were applied. Twenty prediction model studies were evaluated using the PROBAST (Prediction Model Risk of Bias Assessment Tool) [], which is specifically designed for evaluating prediction model studies. PROBAST assesses bias across 4 domains: participants, predictors, outcomes, and analysis. Five observational studies were evaluated using the Newcastle-Ottawa Scale [], while the only randomized controlled trial (RCT) intervention study was assessed with RoB 2 (Cochrane Collaboration) [].


ResultsStudy Selection

The study selection process is illustrated in the PRISMA flow diagram ().

Figure 2. PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) flow diagram.

The systematic search yielded a total of 1683 records: 125 from PubMed, 155 from IEEE Xplore, 922 from Scopus, and 481 from Embase. Prior to screening, 528 duplicate records were removed, resulting in 1155 unique records for title and abstract screening. Based on the predefined inclusion and exclusion criteria, 1098 records were excluded at this stage.

The remaining 57 articles were retrieved for full-text assessment. Of these, 2 records could not be accessed despite efforts to contact the authors via ResearchGate. As a result, 55 full-text articles were assessed for eligibility.

Full-text assessment followed the same inclusion and exclusion criteria as title and abstract screening. After full-text screening, 26 studies [-] met the inclusion criteria and were included in this systematic review.

Study CharacteristicsExacerbation Prediction

The key characteristics of the studies related to exacerbation prediction are summarized in . Across the 20 included studies [-], sample sizes range from 9 to 177 participants, and study durations vary from 1 month to 4 years. Some studies are outliers. For example, Wu et al [] do not explicitly state the length of the monitoring period, and in the study by Lin et al [], the follow-up of the participants ranged from 16 to 448 days, with an average testing period of 206 days per patient.

Definitions of exacerbation also varied substantially. In 3 studies [,,], exacerbations were identified by clinical judgment; 4 studies [,,,] used event-based definitions such as hospitalization or changes in medication. Two studies relied on symptom-based criteria [,], 2 adopted the GOLD definition [,], and one relied on self-reported exacerbations, validated by clinicians []. In the study by Wu et al [], the definition was not clearly reported (“NA”), but it is highly probable that they adopted the same of their previous study []. In other studies, authors used custom rule-based criteria, often operationalized in the study’s algorithm.

Table 1. Characteristics of included studies (exacerbation prediction).Study (year)Population (study duration)Exacerbation definitionModelParameters monitored and samplingWu et al [] (2021)67 (4 months)GOLD guidelinesML and DNNSymptoms - activity, HR, Env/daily - continuousMerone et al [] (2017)22 (6 months)CliniciansBinary FSMSpO2, HR/3 times a dayAtzeni et al [] (2022)101 (6 months)Self-reported, clinicians validationClustering and MLSymptoms, sleep quality, and PEF, EHR - Env/daily - continuousYin et al [] (2024)66 (6 months)GOLD guidelinesMLRespiratory variables, breathing sounds/2 times a dayFernandez-Granero et al [] (2018)15 (6 months)Symptom-basedDecision Tree ForestBreathing sounds/dailyWu et al [] (2022)177 (NA)NAML and DNNSymptoms - activity, HR, Env/daily - continuousKronborg et al [] (2021)9 (2 years)Event-basedMLSpO2, HR, BP/daily to weeklyFernandez-Granero et al [] (2015)15 (6 months)Event-basedPCA-SVMBreathing sounds/dailyNunavath et al [] (2018)94 (2 years)Medical ruleLSTMSymptoms, HR, SpO2/dailyCrooks et al [] (2021)28 (90 days)Event-basedXGB and rule-based systemCough monitoring/continuous at the bedsideKolozali et al [] (2023)106 (6 months)CliniciansPLCA-LDS-3DAudio and PEF - Env, accelerometer/daily - continuousShah et al [] (2017)110 (1 year)Event-basedShort-period trend analysis and LRHR, SpO2, RR/daily to weeklyOrchard et al [] (2018)135 (1 year)HospitalizationNeural NetSymptoms, medications, HR, SpO2, outdoor weather/dailyVan der Heijden et al [] (2014)10 (4 weeks)Symptom-basedTemporal Bayesian NetworkSymptoms, FEV1, SpO2/dailyMohktar et al [] (2015)21 (1 year)Rule-basedCARTFEV1, SpO2, RR, HR, T, weight/dailyJin et al [] (2018)22 (1 month)HospitalizationMLRespiratory variables, SpO2/4
hours a dayLin et al [] (2020)16 (2 weeks-1 year)ER visitsTrend analysis and rule-based systemActivity, sleep/continuousMoraza et al [] (2025)149 (4 years)Rule-basedCatBoost and DLSymptoms, HR, RR, SpO2, steps, T/dailyPatel et al [] (2021)90 (6 months)CliniciansDecision TreeSymptoms - FEV1 - CRP daily
- weekly - at alerting thresholdsHsiao et al [] (2023)13 (6 months)Rule-basedRFHRV/3 times a day

aGOLD: Global Initiative for Chronic Obstructive Lung Disease.

bML: machine learning.

cDNN: deep neural network.

dHR: heart rate.

eEnv: environmental variables.

fFSM: Finite State Machine.

gSpO2: peripheral capillary oxygen saturation.

hPEF: peak expiratory flow.

iEHR: electronic health records.

jNot available.

kBP: blood pressure.

lPCA-SVM: principal component analysis – support vector machine.

mLSTM: Long Short-Term Memory.

nXGB: Extreme Gradient Boosting.

oPLCA-LDS-3D: Probabilistic Latent Component Analysis – Linear Dynamical System in 3D.

pLR: logistic regression.

qRR: respiratory rate.

rCART: Classification and Regression Trees.

sFEV1: Forced Expiratory Volume in 1 second.

tT: temperature.

uER: emergency room.

vDL: deep learning.

wCRP: C-reactive protein.

xRF: random forest.

yHRV: heart rate variability.

Most of the included studies used ML techniques, with a few also incorporating deep learning models. Notably, time-series analysis was explicitly used in only a subset of the studies, whereas many applied ML techniques to static or aggregated data. Of the reviewed studies, 13 reported their best performance using traditional ML techniques [,-,,,-], while 2 relied on rule-based systems [,], 1 on a Bayesian model [], 1 on a finite state machine [], and 3 on deep learning architectures [,,]. Overall, the prevalence of simpler modeling approaches is consistent with the generally small size of the datasets used in these studies.

The monitored variables and sampling frequencies exhibit significant heterogeneity in terms of both sensor modalities and data collection protocols. The variables tracked most frequently are heart rate (HR; 9 studies [,,,,,,,,]), peripheral capillary oxygen saturation (SpO2; 9 studies [,,,,,,-]), and symptoms (8 studies [,,,,,,,]).

To clarify which parameters are being monitored in each study, we have categorized the papers according to the physiological and environmental variables they focus on (1) monitoring of physiological variables (eg, oxygen saturation and HR) [,,,,,,,,], (2) monitoring of respiratory variables (eg, forced vital capacity [FVC] and peak expiratory flow [PEF]), breathing sounds, or coughing [,,,,], (3) monitoring of environmental variables [,], (4) monitoring of both physiological and environmental variables [,,], and (5) monitoring also forced expiratory volume in 1 second (FEV1) [,,,].

The only exceptions stand in Patel et al [], which is monitoring FEV1 and finger-prick C-reactive protein, where the latter is a physiological variable but invasively measured from blood samples.

The sampling frequency varied across studies. Daily sampling was used in 7 studies [,,,,,,], while a few studies collected measurements multiple times per day (2‐3 times). Some studies collected data weekly, and 6 studies [,,,,,] used continuous monitoring. Only in the study by Lin et al [], the sampling frequency explicitly stated, set at 20 Hz. For the remaining studies, no raw sampling is specified. Specifically, in Wu et al [,], data synchronization occurs every 15 minutes, while in the study by Atzeni et al [], it happens every 30 hours. Once again, in the study of Crooks et al [], the system continuously analyzes audio input, but the actual sampling frequency of the raw audio is not specified. In the study by Kolozali et al [], the sampling frequency is not explicitly reported. However, based on the data volume (54 million points over 106 participants), acquisition intervals likely range from subminute to several minutes depending on the sensor.

As the review focuses on wearables and IoT sensors, it is important to note that only a minority of the studies explicitly mention the specific devices used, and that there is substantial heterogeneity in device types, specifically wearable and portable sensors.

Wearable devices, which are continuously worn on the body, included the Fitbit Versa smartwatch and EDIMAX AirBox environmental sensor [,], custom-made Personal Ambient Monitor worn around the waist [], a stationary microphone (Jabra 510+, Jabra GN) at the bed table [], air pollution sensors and activity-related sensors, not further specified [], accelerometer-based wrist-worn wearable devices: GeneActiv [], O2 of Zoetek Inc smartwatch [].

Portable devices, which are intermittently used or carried by the patient, included pulse oximeter ([,,] which included also a pedometer), portable spirometers and electronic stethoscopes [], electronic sensor ad hoc designated [], electronic stethoscopes and a microphone [], home monitoring unit (TeleMed-Care Health Monitor: TMC-Home, TeleMedCare Pty Ltd) [], noninvasive ventilator (RESmart GII BPAP Y-25T, BMC Medical Co, Ltd) [], and Vitalograph, Ireland, to measure FEV1 [].

This taxonomy highlights the lack of standardization in both device selection and reporting across studies, complicating direct comparison and synthesis of findings.

Variables Changes in Prodromal Phases

The key characteristics of the studies related to the analysis of physiological variables in prodromal phases are presented in .

The 6 included studies [-] report sample sizes ranging from 16 to 104 participants, with durations spanning 21 days to 6 months. Two studies stand out as outliers: in Kim et al [], it is not explicitly stated, and in Cooper et al [], all participants collectively provided 2618 days.

The definitions of exacerbation also varied. Two studies used event-based definitions [,], 2 studies relied on symptom-based criteria [,], and one adopted multiple definitions [], in order to see how many different exacerbations were captured, according to the applied definition. In one instance, the definition was not clearly reported (“NA”), not even in the study protocol [].

Table 2. Characteristics of included studies (variable changes in prodromal phases).StudyPopulation (study duration)Exacerbation definitionModelParameters monitored and samplingHawthorne et al [] (2022)35 (6 weeks)Symptom-based (mild), symptom-based and medication (moderate), hospitalization (severe)NARR, HR, T, activity/continuousCoutu et al [] (2024)21 (21 days)Symptom-basedLinear mixed-
effects modelsRR, RRV, HR, HRV, T, ED activity, activity, sleep/continuousZimmerman et al [] (2020)16 (8‐9 monthsIncrease in respiratory symptoms requiring oral corticosteroids and/or antibiotics, with or without medical review and/or hospitalization.Linear mixed-
effects modelsRespiratory variables/dailyAl Rajeh et al [] (2020)83 (6 months)Event-basedNASymptoms - PEF, HR, SpO2/daily- daily vs
overnightCooper et al [] (2020)17 (2618 patient-days)MultipleSPC,
concordance analysisSymptoms, respiratory variables, SpO2/dailyKim et al [] (2021)104 (NA)NALRPM2.5/continuous

aNot available.

bRR: respiratory rate.

cHR: heart rate.

dT: temperature.

eRRV: respiratory rate variability.

fHRV: heart rate variability.

gED: electrodermal activity.

hPEF: peak expiratory flow.

iSPO2: peripheral capillary oxygen saturation.

jSPC: statistical process control.

kLR: logistic regression.

lPM2.5: particulate matter ≤2.5 µm.

The included studies used statistical analysis to inspect the correlations between exacerbations and variation in physiological variables, including HR, respiratory variables, and in one study [], PM2.5 (particulate matter ≤2.5 µm) is considered. Three studies collect data continuously [,,], 2 daily [,], and one assesses the difference between once daily and overnight monitoring [].

Regarding the devices used, it can be highlighted that 4 studies [,,,] out of 6 [-] use wearable devices, specifically:

Hawthorne et al []: chest-worn multiparameter monitoring system (Equivital LifeMonitor EQ02)Coutu et al []: a biometric wristband (EmbracePlus, Empatica Inc) and a biometric ring (Oura Gen III, Oura Health Oy)Al Rajeh et al []: wristband pulse oximeter (Nonin 3150) for overnight monitoring (measurement recorded every 4 s), or finger pulse oximeter (Nonin G92) for once-daily (morning) monitoring.Kim et al []: IoT to measure PM2.5 indoor: a sensor-based light scattering measurement device (CP-16-A5, Aircok Inc).

Meanwhile, Zimmermann et al [] and Cooper et al [] both use portable devices to monitor respiratory trends daily. Specifically, the former uses a FOT (Forced Oscillation Technique) telemonitoring device (Resmon Pro Diary, Restech Srl) to monitor respiratory variables every morning, and the latter RPM (handheld spirometer [SpiroPro, eResearch Technology] and pulse oximeter [Onyx II, Nonin Medical]).

Risk of Bias

The only RCT [] has been assessed as low risk in the randomization process, measurement of the outcome, and selection of the reported results. However, certain issues have been identified regarding deviations from the established interventions. These include the unfeasibility of binding the participants and a high rate of participant dropout. Due to these concerns, a high risk has been assigned to missing data.

The Newcastle-Ottawa Scale results are reported in . Overall, no study raises concerns.

Table 3. Risk of bias assessment (Newcastle-Ottawa Scale).Study (year)Selection (0‐4)Comparability
(0‐2)Outcome (0‐3)Total score (0‐9)Hawthorne et al [] (2022)3238/9Coutu et al [] (2024)3238/9Zimmerman et al [] (2020)3238/9Cooper et al [] (2020)3238/9Kim et al [] (2021)3227/9

Finally, most of the studies are related to prediction models; therefore, the risk of bias has been assessed with the PROBAST, which evaluates Participants and Data Sources (P), Predictors (Pr), Outcome (O), and Analysis (A). The results are summarized in . A high risk of bias was identified in the study by Wu et al [] due to an unclear definition of outcomes and ambiguity in the sample breakdown, which adds uncertainty to the findings. In the study by Kronborg et al [], the use of a very small cohort (9 patients selected from 108), a low number of positive samples (17 exacerbation periods), and reliance on self-reported data introduce a high risk of bias. Nunavath et al [] present potential bias arising from synthetic oversampling, exclusion of one class, and the use of triage labels as proxies for exacerbation severity. Kolozali et al [] apply complex modeling techniques without external validation and with unclear measures to control for overfitting, raising concerns about model reliability. In the study by van der Heijden et al [], the small sample size, the use of self-reported predictors that are not clearly linked to outcomes, and the use of symptoms to both define exacerbations and as input features for prediction models contribute to a heightened risk of bias.

In the “Discussion” section, we discuss how biases in study design and analysis can lead to inflated predictive performance metrics.

Table 4. Risk of bias assessment (PROBAST).Study (year)ParticipantsPredictorsOutcomeAnalysisWu et al [] (2021)LowLowLowSome concernsMerone et al [] (2017)LowLowLowSome concernsAtzeni et al [] (2022)LowLowLowLowYin et al [] (2024)LowLowUnclearLowFernandez-Granero et al [] (2018)LowLowLowSome concernsWu et al [] (2022)LowLowHighUnclearKronborg et al [] (2021)HighLowSome concernsHighFernandez-Granero et al [] (2015)LowLowLowSome concernsNunavath et al [] (2018)HighLowSome concernsHighCrooks et al [] (2021)LowLowLowLowKolozali et al [] (2023)LowLowLowHighShah et al [] (2017)LowLowLowLowOrchard et al [] (2018)LowLowLowLowVan der Heijden et al [] (2014)LowSome concernsHighSome concernsMohktar et al [] (2015)LowLowLowLowJin et al [] (2018)LowLowLowSome concernsLin et al [] (2020)LowLowLowSome concernsMoraza et al [] (2025)LowSome concernsLowSome concernsPatel et al [] (2021)LowLowLowLowHsiao et al [] (2023)LowLowLowSome concerns

aPROBCAST: Prediction model Risk of Bias Assessment Tool.

Study AnalysesExacerbation Prediction

The results of the studies related to exacerbation prediction are summarized in . As anticipated, a broad range of ML and rule-based models have been deployed, with most aiming to forecast exacerbation of COPD between 1 and 7 days ahead. Furthermore, several studies reported overall accuracy without sufficient sensitivity-specificity trade-offs, thereby limiting assessment of clinical utility, and a lack of standardized reporting metrics hinders cross-study comparability. For this reason, we focus our comparisons on studies reporting the same performance metrics.

Table 5. Results of included studies (exacerbation prediction).Study (year)OutcomesPerformance metricsNotesWu et al [] (2021)Early detection of exacerbations (7
days)Accuracy=0.9357, AUROC=0.9699, sensitivity=0.9452, specificity=0.9253, precision=0.9393, F1-score=0.9323DNN on balanced dataset+feature importanceMerone et al [] (2017)Detection of the onset of ECOPDAccuracy=98.4, recall=92.9, specificity=99.3, precision=95.1, F1-score=94.0Participant specificAtzeni et al [] (2022)Early detection of
ECOPD (1
day)Cluster 1: AUC=0.90, AUPRC=0.7, sensitivity=0.83, specificity=0.86; Cluster 2: AUC=0.82, AUPRC=0.56, sensitivity=0.75, specificity=0.78KNN+RF
+ SHAPYin et al [] (2024)Predict ECOPD (k=1, 2, 3 days)AUC (k=1)=0.9721 (95% CI 0.9623‐0.9810), F1-score=0.846CatBoost, balanced test set+feature importanceFernandez-Granero et al [] (2018)Early detection of ECOPD (4.4 days margin)Accuracy=87.8%, sensitivity=78.1%, specificity=95.9%, PPV=94.1%, NPV=83.9%, F1-score=0.8NAWu et al [] (2022)Early detection of ECOPD (7 days)Accuracy=72.4%, sensitivity=62.5%, specificity=78.3%, precision=63.3%, F1-score=62.9%DNN,
external test set, feature
importance, and SHAPKronborg et al [] (2021)Classification of prodromal periods (14 days)AUC=0.95, sensitivity=0.94 (AUC gain=0.24 in a 2-layer model)SVM (RBF), 2-layer modelFernandez-Granero et al [] (2015)Early detection of ECOPD (mean 5, SD 1.9 days)Accuracy=75.8%, sensitivity=73.76%, specificity=97.67%, PPV=84.66%, NPV=95.53%NANunavath et al [] (2018)ECOPD prediction (k=1, 3 days)1-day: accuracy=82.5%, loss=0.50; 3-day: accuracy=88%, loss=0.345Excludes moderate risk patientsCrooks et al [] (2021)Alarm system (3 days ahead)Cough: alerts 3.4 (SD
2.8) days early (success rate=45%), 1 false/100 days; Questionnaire: alerts 3.4 (SD 2.9) days early (success rate: 88%), 1 false/10 daysParticipant-specific baseline for cough frequencyKolozali et al [] (2023)Prediction of daily symptoms (1 day in advance)Accuracy=38.67% and F1-score=0.19 (personalized), 19% and 0.10 (population)Includes recovery phaseShah et al [] (2017)Early detection of ECOPD (7 days)Mean AUC=0.682 (95% CI 0.681‐0.682),
specificity=36% (at 80% sensitivity), 68% (at 60% sensitivity)Analysis of prodromal variablesOrchard et al [] (2018)Next day hospital admission predictionAUC=0.740 (95% CI 0.673‐0.803)No improvement with weather dataVan der Heijden et al [] (2014)Early detection of ECOPD (1 day)Dval AUC=0.90, TPR=0.87, FPR=0.12; Ddedup AUC=0.82, TPR=0.78, FPR=0.19External validation
cohort, Bootstrapping datasetMohktar et al [] (2015)Classify high or low risk (1-day advance)Accuracy=71.8%, specificity=80.4%, sensitivity=61.1%, PPV=71.4%, NPV=72%Stable days excluded,
feature importanceJin et al [] (2018)Classification of prodromal periods (7 days)Accuracy=74.5% (LDA), 73.7% (SVM), 75% (RF); sensitivity=NA, 77.6%, 78.3%; specificity: lowest, 42.9%, 42%Feature importanceLin et al [] (2020)30-day hospital readmission predictionSensitivity=62.96%, precision=37.78%,
miss rate=37.04%, FDR=62.22%NAMoraza et al [] (2025)Early detection of ECOPD (3 days)Test AUROC=0.91, AUPRC=0.53; prospective: AUROC=0.89, AUPRC=0.56SHAP analysis, prospective validation dataPatel et al [] (2021)Early detection of ECOPD (7, 3 days [median, IQR])Sensitivity=97.9%, specificity=84%,
PPV=38.4%, NPV=99.8%, accuracy=85.3%Participant-specific baselineHsiao et al [] (2023)Early detection of ECOPD (7 days)Accuracy=0.94%, precision=0.57%, sensitivity=0.48%, specificity=0.93%Rule-based predictions

aAUROC: area under the receiver operating characteristic curve.

bDNN: deep neural network.

cECOPD: exacerbation of chronic obstructive pulmonary disease.

dAUC: area under the curve.

eAUPRC: area under the precision-recall curve.

fKNN: k-nearest neighbors.

gRF: random forest.

hSHAP: Shapley Additive Explanations.

iPPV: positive predictive value.

jNPV: negative predictive value.

kNA: not available.

lSVM: support vector machine.

mRBF: radial basis function.

nDval: validation dataset.

oTPR: True positive rate.

pFPR: false positive rate.

qDdedup: deduplicated dataset.

rLDA: linear discriminant analysis,

sFDR: false discovery rate.

Deep neural networks and tree-based ensemble methods achieved strong performance with AUC values approaching 90%. Wu et al [] reported one of the best-performing models using a deep neural network with a 7-day prediction window, achieving an area under the receiver operating characteristic curve (AUROC) of 0.97, accuracy of 93.6%, and sensitivity of 94.5%. Nonetheless, the test set is balanced 1:1, thus not representing the real-world distribution. Similarly, CatBoost-based models by Yin et al [] and Moraza et al [] achieved AUROC values above 0.89, indicating strong discriminatory power for short-term prediction. However, Yin et al [] used a 5-fold cross-validation with a balanced test set, which does not reflect real-world class imbalance.

Other models, such as those by Atzeni et al [] and Kronborg et al [], also achieved high AUROC scores (≥0.90) in specific patient clusters or multilayer support vector machine setups.

Only Wu et al [] and van der Heijden et al [] performed external validation on a different cohort, while Moraza et al [] used prospective validation data sampled from the sample cohort.

Studies using random forest models (eg, [,,]) demonstrated varying levels of performance. Wu and Jin typically reported moderate overall accuracies between 72% and 80%, with trade-offs in sensitivity and specificity. Jin et al [] reported relatively balanced sensitivity (~76%‐78%) but consistently low specificity (~41%‐43%), suggesting a higher false positive rate. In contrast, Hsiao et al [] reported a notably high accuracy of 0.94 and specificity of 0.93, but with lower precision (0.57) and sensitivity (0.48), highlighting strong negative class identification at the cost of reduced detection of true positives.

Models with rule-based or alert systems (eg, []) demonstrated real-world feasibility with fewer false positives, though at the expense of lower sensitivity (45% ECOPD detection with cough-based alarms), also using baseline personalization. Personalization appeared to significantly improve predictive performance. Patel et al [] used a personalized baseline approach, achieving high sensitivity (97.9%) and negative predictive value (99.8%), despite a modest positive predictive value (PPV; 38.4%), indicative of a low false-negative rate, which is important for safety but less efficient for resource use. Other studies have also underscored the value of patient-specific modeling (eg, [] with 98.4% accuracy in a participant-specific setup).

Lower-performing models (eg, []) demonstrated limited predictive power in both personalized (38.7%) and population-level (19%) settings, potentially due to inclusion of recovery phases or inadequate feature representation. Similarly, Lin et al [] achieved moderate sensitivity (62.9%) but low precision (~38%), leading to high false discovery rates.

Finally, 3 studies [,,] highlighted temporal margins for predictions: they reported early detection windows of ~3‐5 days with variable sensitivity and precision, balancing warning time with model reliability. Notably, Nunavath et al [] excluded deteriorating patients from their model, considering only stable and urgent cases—once again failing to reflect real-world distributions.

Studies can be further categorized according to the selection of the prediction window [,,,,]. Predict exacerbations 7 days ahead. Choosing a shorter time window (1‐3 days), there are [,,-,,,]. Meanwhile, the other papers focus on classification tasks. As shown in , high performance is reported across both short and long prediction windows.

Some studies analyzed which variables were most important for prediction. Wu et al [] highlighted COPD Assessment Test score, fine particulate matter (PM10), and carbon monoxide; Atzeni et al [] emphasized symptoms in previous days and a higher frequency of exacerbations in the past 30 days; and Mohktar et al [] identified raw FEV1 and mean SpO2 as key features. Shah et al [] observed lower SpO2, higher respiratory rate (RR), and elevated HR during the prodromal phase, concluding that SpO2 was the most predictive feature. Overall, all studies, apart from Orchard et al [], concluded that including all features led to better results.

Finally, models often lack explainability, which can be a fundamental element to build clinicians’ trust in the black box model. In this review, interpretability was addressed in only 7 [,,,,-] out of 20 studies [-].

Variables Changes in Prodromal Phases

The results of the studies related to the analysis of changes in physiological variables in prodromal phases are summarized in . Hawthorne et al [] identified significant changes in HR and RR up to 3 days prior to exacerbations, with increases of 8.1 bpm and 2.0 breaths/minute, respectively. Coutu et al [] found significant associations between higher EXACT-PRO (Exacerbations of Chronic Pulmonary Disease Tool–Patient-Reported Outcome) scores and reduced RR variability, step count, and sleep efficiency, suggesting that deterioration in these metrics may indicate worsening symptoms. Similarly, Zimmermann et al [] demonstrated that variability in inspiratory reactance was significantly related to COPD Assessment Test scores.

Table 6. Results of included studies (variable changes in prodromal phases).Study (year)OutcomesPerformance metricsHawthorne et al [] (2022)Monitor changes in vital signs 3, 2, and 1 days before ECOPDHR and PA were associated with EXACT score (P<.001). Three days prior to exacerbation: RR increased by 2.0 (SD 0.2) breaths/min (7 of 11 cases), HR increased by 8.1 (SD 0.7) bpm (9 of 11)Coutu et al [] (2024)Estimate the association between physiological variables and EXACT-PRO scoreSignificant unadjusted associations included RR variability (−1.45 [−2.84, −0.073] points per breath/min), daily step count (−0.56 [−0.82, −0.31] points per 1000 steps), and sleep efficiency (−0.12 [−0.20, −0.037] points per % asleep), with bracketed values representing 95% CIs.Zimmerman et al [] (2020)Timing of variability in FOT measures and symptom change
before AECOPDXinsp and SDXinsp were related to mean CAT score (0.59 [1.02, 0.15], P=.009; 1.57 [0.65‐2.49], P=.001)Al Rajeh et al [] (2020)Compare overnight versus daily HR and SpO2 monitoringDaily HR: max increase +7 bpm on day 1 (P=.007, not clinically significant). Overnight HR: +10 bpm on day 1 (P=.04). Composite oximetry score increased from day 7 to 0 (overnight), and on days 1 and 0 (daily): sensitivity 84.6%, PPV 91.7%Cooper et al [] (2020)Identify monitored variables linked to exacerbationsAgreement between FVC drop (7-day avg 1.645 SD) and self-reported health care use (Cohen κ=0.747, P<.001); bronchodilator use and Anthonisen criteria (Cohen κ=0.611, P<.001)Kim et al [] (2021)Association between PM2.5 and questionnaire responsesMean PM2.5 over 4 months was higher in patients with exacerbations (22.89, SD 5.52 g/m3) vs no-exacerbation patients (20.36, SD 4.63 g/m3), P<.02

aECOPD: exacerbation of chronic obstructive pulmonary disease.

Comments (0)

No login
gif