Body temperature is a crucial physiological measure, reflecting the fundamental homeostatic process of thermoregulation in both healthy individuals and in response to disease. In clinical practice, abnormalities in core body temperature such as fever or hypothermia, are used as primary indicators of infection, inflammation, or systemic compromise (Zhao and Zhang 2024). However, while core body temperature is tightly regulated within a narrow range, peripheral skin temperature is inherently dynamic. This variability reflects a complex interaction between heat loss (e.g. through vasodilatation) and heat gain (e.g. through metabolic activity). Pattern analysis of these fluctuations, termed skin temperature variability (STV), serves as an umbrella term encompassing multiple analytical approaches, including total variability (standard deviation (SD)), short-term and long-term variability (Poincaré plot), complexity (entropy, fractal measures) and long-term circadian amplitude. Together, these metrics may provide a unique window into the integrity of autonomic, vascular, and thermoregulatory control systems (Bottaro et al 2020).
Emerging evidence shows that changes in STV dynamics, across both short-term (minute-to-minute) and long-term (circadian) fluctuations, signal systemic dysfunction. Computational approaches such as entropy and detrended fluctuation analysis (DFA) revealed that reduced temperature complexity is associated with higher mortality in critically ill patients, including those with sepsis and multiple organ failure (Varela et al 2006, Papaioannou et al 2013). This diminished complexity signifies a loss of physiological flexibility and adaptive capacity, indicative of control system failure. Similarly, variability metrics derived from proximal skin temperature recordings in patients with cirrhosis have been shown to predict survival independently of conventional severity scores (Bottaro et al 2020, Abid et al 2024). For chronic conditions, large population studies suggest that a blunted circadian amplitude of wrist temperature is associated with increased risk of cardiometabolic and inflammatory diseases, such as diabetes, non-alcoholic fatty liver disease (NAFLD or MASLD), and renal failure (Brooks et al 2023). We hypothesise that altered STV profiles are associated with adverse outcomes, however, the direction and meaning of these alterations may depend on disease context, measurement timescale, and analytical method.
The rationale for studying STV mirrors the established field of heart rate variability (HRV), where a loss of complexity signals impaired resilience and adverse outcomes (Mani et al 2009, Shaffer and Ginsberg 2017, Bhogal et al 2019, Abid and Mani 2022). However, STV is distinct from HRV; while HRV primarily reflects autonomic modulation of the cardiac rhythm, STV specifically integrates peripheral vascular reactivity, sweat gland activity, and metabolic heat production. Thus, STV may provide a more comprehensive measure of physiological control and thermoregulatory coupling that HRV cannot capture. In addition, conventional HRV measures cannot be used to assess autonomic function in individuals with atrial fibrillation or other types of cardiac arrhythmias (Oyelade et al 2020), and they are subject to a high degree of noise due to movement or displacement of sensors, whereas STV is not directly affected by arrhythmias and is less susceptible to noise during long-term measurement.
Despite the potential of STV, the field is currently limited by a lack of standardisation. Studies vary widely in measurement sites (e.g. wrist vs. hypochondrium) and sampling frequencies (seconds vs. hours) which complicates the translation of these findings into routine clinical practice. If a physiological marker is associated with prognosis, this suggests that it is linked to essential physiological mechanisms important for survival. Therefore, it is important to systematically evaluate whether a physiological marker such as STV is associated with prognosis. To the best of our knowledge, unlike HRV, STV has not yet been systematically evaluated to establish its comparative prognostic utility across different medical conditions.
Thus, the aims of this systematic review are to:
(1)
Synthesise current evidence on the prognostic role of various STV metrics in human studies across different populations and health states.
(2)
Characterise the technical heterogeneity in measurement sites and analytic approaches.
(3)
Evaluate the associations between STV and clinical outcomes, defined broadly to include mortality risk, disease severity indicators, and future disease risk stratification.
(4)
Identify gaps and propose future research directions, including the potential integrations of STV monitoring into routine clinical practice.
2.1. Overview and ethics statementThis systematic review was conducted in accordance with the PRISMA 2020 guideline (Page et al 2021). The review synthesises current evidence on the prognostic value of STV in diverse clinical settings, incorporating studies inception to 5th June 2026. This systematic review extracted only previously published data and did not involve any new data collection from human participants or animals. Thus, no ethical approval was required.
2.2. Eligibility criteriaInclusion criteria were categorised using the SPIDER framework (Sample, Phenomenon of Interest, Design, Evaluation, Research type), shown in table 1. Studies were eligible for inclusion if the following criteria were met:
(1)
Human participants;
(2)
STV was dynamically measured (e.g. short-term fluctuations, circadian patterns, and entropy) and linked to prognostic outcomes, with ⩾3 time points in the temperature series;
(3)
original research articles, including randomised controlled trials (RCTs), prospective or retrospective observational studies, cohort studies, or case-control studies
Table 1. The SPIDER tool applied to the review questions.
Human participants of any age where STV was assessed in relation to health or clinical outcomesii.
Phenomenon of Interest
STV, including short-term fluctuations, circadian patterns, or entropy/complexity metrics Published peer-reviewed literature of any primary research design Clinical outcomes, including prognosis, mortality, morbidity, organ dysfunction, infection detection, or risk stratification Quantitative or mixed methods peer-reviewed primary studiesNo restrictions were placed on age, gender, medical condition, or healthcare setting. Non-English studies were excluded. Studies using surrogate outcomes unrelated to prognosis (e.g. thermoregulation mechanisms only) or lacking measurable clinical endpoints were excluded. The following literature types were excluded: conference abstracts, reviews, protocols, grey literature (e.g. dissertations), and animal studies.
2.3. Search strategyA comprehensive literature search was conducted using Ovid MEDLINE, EMBASE, AMED from database inception to 5th June 2026. The search strategy used a combination of subject headings, free-text terms and field tags (covering title, keyword, and keyword plus fields). The complete search syntax for all databases is provided in Supplementary material S1. As detailed in table 2, the strategy was organised around the intersection of three conceptual blocks: [Skin temperature] AND [Variability measure] AND [Outcome]. Manual reference checking and forward and backward citation tracking were performed to identify additional relevant studies. A filter for human studies were applied. The search included peer-reviewed primary research articles, with no restriction to titles or abstracts, to ensure a broad and inclusive approach to literature retrieval.
Table 2. Search terms used to identify papers related to skin temperature variability and outcome (Ovid).
Skin temperatureVariability measureOutcomeSkin temperatureVariability analysisPrognosisCutaneous temperatureEntropy (sample entropy, approximate entropy, multiscale entropy, fuzzy entropy, permutation entropy)SurvivalPeripheral temperatureFractal dimension, detrended fluctuation analysis,MortalityWrist temperaturePoincaré plotMorbidityBody temperatureNonlinear dynamics,Risk Assessment, deterioration2.4. Study selectionStudy selection followed a two-step screening process. Titles and abstracts were initially screened for relevance, followed by full-text reviews to confirm eligibility. Any disagreements were resolved through discussion between two independent reviewers (Eubi Chan [EC] and Angelica Blotto [AB]), and any discrepancies were discussed with Ali R Mani [AM] to reach a consensus. All reasons for exclusion were documented, and the selection process is shown in the PRISMA flow diagram (figure 1).
Figure 1. PRISMA flow-chart of study selection.
Download figure:
Standard image High-resolution image 2.5. Data extraction and managementData were extracted from the included studies independently by two reviewers (EC and AB) using a standardised extraction template. Any discrepancies were resolved through discussion and consensus with a third reviewer (AM). To ensure a comprehensive synthesis, the following data items were extracted as shown in tables 3 and 4:
(i)
Characteristics of all included studies: Author, population/condition, country, study type, sample size, measurement site, sampling frequency, specific STV metrics, and key finding(s).
(ii)
Predictive efficacy and effect sizes: Primary outcome (e.g. mortality, incident disease), follow-up duration, effect size/performance (including hazard ratios [HR], odds ratios [OR], or area under the curve [AUC]), and adjustment variables/benchmarking.
Table 3. Characteristics of all included studies (n = 10). Abbreviations: AKI = Acute Kidney Injury; ApEn = Approximate entropy; CRRT = Continuous Renal Replacement Therapy; CV = Coefficient of Variation; IHD = Ischaemic Heart Disease; SD = Standard Deviation; SIRS = Systemic Inflammatory Response Syndrome; SOFA = Sequential Organ Failure Assessment Score; STV = Skin Temperature Variability.
Author (Year)Population/ conditionCountryStudy typeSample sizeMeasurement siteSampling frequencySTV metricKey finding (S)Abid et al (2024)Cirrhosis & SepsisUK, ItalyProspective observational40 (30 men and 10 women)Hypochondrium, Clavicle, ThighEvery 3 minExtended Poincaré plots (SD1, SD2)1-hour proximal skin temperature recordings (daytime) predicted 12 month survival independent of severity of liver disease. SD2 from 1-h recordings was prognostic; longer durations enhanced prediction but 1-h recordings were sufficient for clinical utility.Pornsirirat et al (2024)AKI on CRRTThailandRetrospective observational300 (176 men and 124 women) Age range: 53-81Axillary & TympanicEvery 2-4 hCV (SD/Mean) & AmplitudeHypothermia occurred in 23.7% of patients within 24 h of CRRT initiation. Older age, lower body weight, higher fluid balance, and higher CRRT dose predicted hypothermia. Increased temperature variability (CV) was associated with higher ICU mortality.Brooks et al (2023)General populationUSA (based on UK biobank data)Prospective observational population-based study91 462Dominant wrist1 sample every 75 s (0.013 Hz)Cosinor fitsAmong 425 disease conditions compared to controls, a total of 73 (17%) disease phenotypes were significantly associated with decreased amplitudes of wrist temperature variability. Overall, subjects with lower amplitude had increased risk of all-cause Mortality. Wrist temperature variability also had large associations with risk of hypertension and in disorders of lipid metabolism, consistent with the link to metabolic syndrome.Bottaro et al (2020)Liver CirrhosisUK, ItalyProspective observational38 (30 men and 8 women) Mean age: 64.62 years ± 10.4 yearsHypochondrium, Clavicle, ThighEvery 3 minExtended Poincaré plots (SD1, SD2), Sample Entropy, DFAReduced short-term temperature variability (SD1) and prolonged memory length were independent predictors of 12 month mortality, independently from severity of liver disease. Absolute temperature was not predictive of outcomes.Amiya et al (2013)IHD ± DiabetesJapanProspective observational66 (43 men and 23 women) Mean age: 68.6 ± 7.3 yearsAxillary3 times per daySD & Amplitude (Max–Min)In non-diabetic patients, body temperature variability correlated with endothelial function as evaluated on flow-mediated dilation. In diabetic patients, this correlation was lost, but higher body temperature variability predicted shorter event-free survival.Papaioannou et al (2013)SepsisGreeceProspective observational21 (13 men and 8 women) Mean age: 62.56 ± 6.6 yearsRight hypochondriumEvery 10 sMultiscale EntropyBoth Shannon and Tsallis entropies were significantly reduced in non-survivors. Entropy metrics predicted ICU mortality better than SOFA score, particularly at very low-frequency scales. Combining entropy with SOFA improved prediction accuracy.Papaioannou et al (2012)ICU (SIRS vs. sepsis vs.septic shock)GreeceProspective observational22 (14 men and 8 women) Mean age: 60.86 ± 5.35 yearsRight hypochondriumEvery 10 sWavelet Multiscale Entropy & Sample entropy (SampEn)Temperature complexity (entropy measures) was significantly reduced in patients with sepsis and septic shock compared to SIRS. Wavelet and multiscale entropy were able to distinguish severity groups, potentially reflecting altered thermoregulation and disease severity.Cuesta et al (2007)Critically ill patients at ICUSpainProspective, observational36 (No sex information reported) Mean age: 63.0 ± 11.7 yearsHypochondriumEvery 10 minApproximate entropy (ApEn)ApEn of temperature time-series is reduced in non-survivors. ApEn can be used as a classifier to predict prognosis at ICU.Varela et al (2006)Critically ill patients at ICUSpainProspective, observational50 25 women (mean age 62.4 yrs) 25 men (mean age 60.8 yrs)HypochondriumEvery 10 minApEn & Detrended Fluctuation Analysis (DFA)Low levels of complexity in the temperature curve were indicators of poor prognosis in patients with multiple organ failureVarela et al (2005)Critically ill patients at ICUSpainProspective, observational24 13 women (mean age 62.4 yrs) 11 Mean (mean age 58.0)HypochondriumEvery 10 minApEnMean temperature ApEn was significantly lower in patients who died than in critically ill patients who survived ICU admission. Reduced complexity of temperature time-series was associated with reduced survival.Table 4. Predictive efficacy and effect sizes of skin temperature variability metrics by primary outcome, follow-up duration, and multivariable adjustment. Abbreviations: AUC = Area Under the Curve in ROC curve analysis; HTN = Hypertension; HR = Hazard Ratio; MELD = Model for End-Stage Liver Disease; OR = Odds Ratio; T2D = Type 2 Diabetes.
Author (Year)Primary OutcomeFollow-Up DurationEffect Size/Performance (95% CI)Adjustment Variables/ BenchmarkingAbid et al (2024)Mortality12 monthsHR: 0.017 (0.000–0.808) for 1-h SD2 in extended Poincaré plot analysis; AUC: 0.66 (0.58–0.74, p < 0.001) in discriminating survivors from non-survivors.Adjusted for MELD score; independent of liver failure.Pornsirirat et al (2024)Mortality30 dAdjusted OR: 1.41 (1.13–1.78, p = 0.003) for CVTempAdjusted for age, weight, fluid balance, and CRRT dose.Brooks et al (2023)All-cause mortality and incident morbidityMedian ∼10 yearsAll-cause mortality: HR: 1.14 (1.04–1.23) Risk of developing morbidities: NAFLD: HR: 1.91 (1.58–2.31)T2D: HR: 1.69 (1.53–1.88)Renal Failure: HR: 1.25 (1.14–1.37)HTN: HR: 1.23 (1.17–1.30)Pneumonia: HR 1.22 (1.11–1.33). Hazard ratios were calculated to assess the risk of all-cause mortality or morbidity associated with a reduction in wrist temperature time-series amplitude below two standard deviations (1.8 °C)Adjusted for age, sex, BMI, ethnicity, and Townsend deprivation index.Bottaro et al (2020)Mortality12 monthsHR: 6.57 × 10−9 (p = 0.021) for SD1 (k = 3) in extended Poincaré plot analysis; AUC: 0.737 (p = 0.011) in discriminating survivors from non-survivors. HR: 1.216 (p = 0.006) for Memory length (−3σ); AUC: 0.691 (p = 0.042) in discriminating survivors from non-survivors.Adjusted for MELD and Child-Pugh scores.Amiya et al (2013)Event-free survival16.4 ± 8.4 monthsEvent-free survival differed significantly between BT SD >0.27 °C and <0.27 °C groups (Kaplan–Meier, log-rank p = 0.012)Correlated with Flow-Mediated Dilation.Papaioannou et al (2013)MortalityICU stayAUC: 0.90 for multiscale entropy of temperature time-series vs. 0.82 for SOFA in discriminating survivors from non-survivorsCombined multiscale entropy + SOFA achieved AUC: 0.94.Papaioannou et al (2012)Disease severityICU stayClassification accuracy >80% for distinguishing SIRS from sepsisDiscriminated between SIRS, sepsis, and septic shock.Cuesta et al (2007)MortalityICU stayAccuracy: 0.72; AUC: 0.73 for distinguishing survivors from non-survivorsNot reportedVarela et al (2006)MortalityICU stayOR: 7.64 (2.38–24.51) per 0.1 decrease in ApEnAge-adjusted using linear regression residuals.Varela et al (2005)MortalityICU stayOR: 15.4- and 18.5 per 0.1 decrease in ApEn minimum and mean respectively The discriminating ability of the mean ApEn value (for a cutoff point of 0.55) offered a sensitivity of 0.91 (95% CI, 0.57–0.99), a specificity of 0.85 (95% CI, 0.54–0.97), a positive predictive power of 0.83 (95% CI, 0.51–0.97), and a negative predictive power of 0.92 (95% CI, 0.60–1.0).Age-adjusted using linear regression residuals.No study authors were contacted for missing data as the published reported provided sufficient detail or the current qualitative synthesis.
2.6. Quality assessmentThe quality of each included study was independently assessed by two reviewers (EC and AB) using the Standard Quality Assessment Criteria for Evaluating Primary Research Papers (QualSyst) for quantitative studies. The QualSyst tool consists of 14 criteria, each scored from 0 to 2, with the option to mark items as ‘not applicable’ (N A). The maximum possible score is 28, with results converted to a percentage. Studies were classified as poor (<50%), fair (50%–69%), good (70%–79%), or strong (>80%) quality (Kmet et al 2004). Discrepancies between reviewers were resolved through discussion until consensus was reached. Supplementary material S2 provides a detailed summary of the quality assessment scores across the included studies.
The methodological quality of the clinical studies was additionally assessed using the Quality in Prognosis Studies (QUIPS) tool (Hayden et al 2013), given the prognostic focus of the included studies. The QUIPS tool evaluates six domains: study participation, study attrition, prognostic factor measurement, outcome measurement, confounding measurement and account, and statistical analysis. Each domain was rated as low, moderate, or high risk of bias. Discrepancies between reviewers were resolved through discussion until consensus was reached. The full domain-level risk of bias assessment is presented in Supplementary material S3.
2.7. Evidence synthesis and development of logic model2.7.1. Rationale and development of the logic modelWe developed a logic model to synthesise the results of the included studies in a meaningful and systematic way, enabling visualisation of relationships between population/condition, temperature measurement site/method, skin temperature metric, and clinical outcome.
Logic models are increasingly recommended in systematic reviews and health technology assessments of complex questions, as they (i) clarify hypothesised causal chains and context–mechanism–outcome links; (ii) identify mediators, moderators, and intermediate outcomes; and (iii) inform subgrouping and sensitivity analyses (Anderson et al 2011, Rohwer et al 2017, Rehfuess et al 2018). Following Rohwer et al (2017), we constructed an iterative, process-based logic model that evolved throughout data extraction and synthesis, allowing incorporation of cumulative evidence and emerging patterns.
As no direct intervention was assessed, we applied a modified PICO framework (i.e. the SPIDER framework mentioned in table 1), representing key components through interconnected boxes for population/condition, temperature measurement site/method, skin temperature metric, and clinical outcomes. Coloured arrows map study-specific pathways, linking populations, measurement approaches, and temperature metrics to prognostic outcomes.
3.1. Overview of study selectionA total of 61 records were identified through database searches, comprising Ovid MEDLINE, Embase, and AMED. After 20 duplicate records were removed, 41 titles and abstracts were screened for eligibility. Of these, 24 records were excluded for not meeting the eligibility criteria, leaving 17 full-text studies for detailed assessment. Nine studies were excluded for the following reasons: Narrative review (n = 1), Core body temperature only, with no skin temperature data (n = 5), No survival-related outcome was evaluated (n = 3). 8 studies met all inclusion criteria. An additional two studies were identified through manual citation tracking, resulting in a total inclusion of 10 studies. A detailed flow of study selection is presented in figure 1.
3.2. Study design and characteristics of included studiesTable 3 summarises the characteristics of the 10 included studies published between 2005 and 2024. These studies represent a geographically diverse range of clinical settings, including the UK, Japan, Greece, Italy, Spain, and Thailand, and target varied patient populations. The evidence base spans acute ICU cohorts with sepsis or multiple organ failure, hospital-based cohorts with cirrhosis and ischaemic heart disease, and specialised critically ill populations such as patients with AKI undergoing continuous renal replacement therapy (CRRT). Additionally, the synthesis incorporates a massive population-based longitudinal cohort from the UK Biobank.
Of the 10 studies, eight utilised a prospective observational design (Varela et al 2005, Varela et al 2006, Cuesta et al 2007, Papaioannou et al 2012, Papaioannou et al 2013, Amiya et al 2013, Bottaro et al 2020), with 1 retrospective cohort study (Pornsirirat et al 2024), and 1 population-based longitudinal study (Brooks et al 2023). Sample sizes varied significantly between study types; clinical cohorts ranged from 21 to 66 participants, while population-level analysis included 91 462 individuals. Measurement protocols also exhibited high variability with sampling frequencies ranging from high-resolution 10-second intervals to sparse manual measurements taken three times daily.
3.3. Assessment of methodological qualityThe methodological quality and risk of bias assessments are summarised in Supplementary Materials S2 and S3. Based on the QualSyst assessment, nine studies were classified as high quality, with total scores ranging from 81.8% to 100%, while one study was classified as being of fair quality (Cuesta et al 2007). No study fell below the quality threshold for inclusion.
The QUIPS assessment showed a generally low to moderate risk of bias for the majority of studies (n = 6), with high risk identified in at least one domain for the remaining four studies. Study attrition was rated as low risk across all included studies (n = 10), and outcome measurement was rated as low risk for 80% (n = 8) of the studies. The primary concerns identified were related to confounding measurement and statistical analysis. Specifically, high risk of bias in analysis was identified for Amiya et al, Cuesta et al and both Papaioannou et al studies (Cuesta et al 2007, Papaioannou et al 2012, 2013, Amiya et al 2013). These ratings were primarily driven by small sample sizes and limited capacity for multivariable adjustment in exploratory diabetic cardiovascular and ICU cohorts.
While most studies utilised high-frequency monitoring, a moderate risk in prognostic factor measurement was noted for Amiya et al (n = 1) due to a sparse sampling rate of three times per day (Amiya et al 2013). Brooks et al was the only study (n = 1; 10%) to be rated as low risk in the analysis domain, benefiting from a large-scale population-based sample size (n = 91 462) (2023). All included studies (n = 10; 100%) were rated as having a moderate risk of bias for study participation, reflecting recruitment from non-population-based clinical cohorts, which may limit representativeness, or insufficient reporting of baseline characteristics.
3.4. Quantitative measures of prognostic performanceThe prognostic efficacy and standardised effect sizes of STV across diverse clinical outcomes are summarised in table 4. For incident disease outcomes, a two-SD decrease (1.8 °C) in wrist temperature amplitude was an independent predictor of multiple conditions over a median follow-up of approximately 10 years, including all-cause mortality (HR 1.14, 95% CI 1.04–1.23), NAFLD (HR 1.91, 95% CI 1.58–2.31), Type 2 diabetes (HR 1.69, 95% CI 1.53–1.88), and renal failure (HR 1.25, 95% CI 1.14–1.37) (Brooks et al 2023). Notably, these associations remained statistically significant in this study after multivariable adjustment for factors such as age, sex, BMI, ethnicity, and Townsend deprivation index.
In acute and critical care settings, STV metrics demonstrated statistically significant predictive power for mortality. Increased axillary temperature variability (CVTemp) was associated with higher ICU mortality in patients receiving CRRT (Adjusted OR: 1.41, 95% CI: 1.13–1.78) (Pornsirirat et al 2024), while a 0.1 decrease in Approximate Entropy (ApEn) was associated with an OR of 7.64 (95% CI: 2.38–24.51) for mortality (Varela et al 2006). The same group of investigators also reported similar findings in a smaller cohort, with ORs of 15.4 and 18.5 for each 0.1 decrease in the minimum and mean ApEn values, respectively, during serial skin temperature recordings (Valera et al 2005) and an AUC of 0.73 in ROC analysis for distinguishing survivors from non-survivors (Cuesta et al 2007). Furthermore, the predictive performance of STV metrics often exceeded or complemented established clinical scores; for instance, combining entropy with SOFA achieved an AUC of 0.94 (Papaioannou et al 2013).
3.5. Skin temperature measurement sites and methodTwo studies measured proximal skin temperature on sites such as the abdomen, intra-clavicular and mid-thigh (Bottaro et al 2020, Abid et al 2024). Five studies recorded skin temperature on the hypochondrium (Varela et al 2005, 2006, Cuesta et al 2007, Papaioannou et al 2012, 2013). Two studies measured axillary temperature (Amiya et al 2013, Pornsirirat et al 2024), with the former additionally measuring tympanic temperature. One study measured peripheral wrist temperature on the dominant wrist (Brooks et al 2023)
Temperature measuring devices varied across studies. Proximal skin temperature measurements on the hypochondrium primarily used thermistor sensor (Datalogger Spectrum 1000; Veriteq Instruments, Richmond, BC, Canada) (Varela et al 2005, 2006, Cuesta et al 2007, Papaioannou et al 2012, 2013). More recently, iButton sensors (model DS1922L-F5, Maxim Integrated, CA, United States of America) have been employed (Bottaro et al 2020, Abid et al 2024). Axillary temperature in Amiya et al (2013) was measured using the Terumo Digital Clinical Thermometer C202, whereas Pornsirirat et al (2024) did not specify the device for axillary and tympanic measurements. Peripheral wrist temperature was measured using an embedded sensor in the Aximetry AX3 wrist-worn device (Brooks et al 2023).
Sampling rate ranged from every 10 s (Papaioannou et al 2012, 2013) to 3 times a day (Amiya et al 2013). Intervention durations were either fixed, for example, three consecutive days (Amiya et al 2013), or outcome-dependent, such as every 10 min from inclusion in the study until discharge from ICU or death (Varela et al 2005, 2006, Cuesta et al 2007). In one study, varying recording lengths (30 min, 1, 2, 3, and 6 h) were sequentially extracted from 24-hour temperature data to determine the shortest duration required to independently predict 12 month survival in patients with cirrhosis (Abid et al 2024). These findings highlight the methodological variability across studies, suggesting a need for standardised and optimised temperature measurement protocols to enhance comparability and clinical applicability.
3.6. Metrics of STVData extraction from the selected studies revealed that no standardised method for measuring STV exists in the literature. Different studies used various analytical approaches depending on the investigators’ perspectives on which aspects of physiological function STV analysis should capture. Overall, some studies assessed overall STV by calculating the SD of the temperature time-series, while others used the Poincaré plot or wavelet transformation to decompose STV into short-term and long-term variability measures. A few studies employed more complex methods to examine patterns of skin temperature fluctuations by assessing entropy, fractal-like structures, or the controllability of time-series dynamics. The motivation behind each metric is summarised below and illustrated in figure 2.
Figure 2. Schematic diagram showing the most common methods used in selected studies for the calculation of skin temperature variability. (A) Standard deviation and amplitude (maximum–minimum) are the most common methods used for measuring temperature variability. (B) In studies where temperature was measured to assess circadian rhythms, the amplitude of skin temperature variability was assessed by the amplitude of fluctuations after cosinor fit of the temperature signal.(C) and (D) The Poincaré plot is used to assess the autocorrelation of temperature time-series. This method allows estimation of short-term (SD1) and long-term (SD2) variability. (E) and (F) Wavelet transformation is used to assess the frequencies of temperature fluctuations at different times. (G) and (H) Entropy analysis (e.g. calculation of Sample Entropy) is used to assess the complexity of the time-series. In this method, entropy is calculated as the negative logarithmic likelihood that a template of size m repeated in the data will also be repeated for a template of size m + 1. (I) and (J) detrended fluctuation analysis (DFA) is used to assess the fractal-like structure of a time-series. In this method, the original time-series is detrended at different scales, and the relationship between the logarithm of the scale and the logarithm of the fluctuation of the detrended signal is assessed. If this relationship is linear, the time-series is fractal-like, and the slope (called the scaling exponent or α) determines the fractal dynamics.
Download figure:
Standard image High-resolution image 3.6.1. SDSD is the square root of the average squared deviation of each temperature value from the mean temperature of the time-series. It quantifies the overall dispersion of temperature values around the mean but does not describe the temporal pattern or sequential structure of fluctuations. This metric was used in studies by Papaioannou et al (2012), Amiya et al (2013), Bottaro et al (2020), and Pornsirirat et al (2024), who also employed the coefficient of variation (CV), a related metric defined as SD divided by the mean. Other central moments of the skin temperature distribution, such as skewness and kurtosis, were not used in the included studies.
3.6.2. Poincaré plot measuresThe Poincaré plot graphically represents the correlation between consecutive points in a time-series (figure 2). It illustrates how each point in the series is influenced by previous points and enables the estimation of short-term and long-term variability by calculating the SD of points perpendicular to the line of identity (SD1) and along the line of identity (SD2), respectively. The physiological interpretation of short- and long-term STV has not been comprehensively studied. It appears that the interpretation of SD1 and SD2 depends on the recording length and sampling rate. For example, in long recordings (e.g. 24 h), SD2 reflects circadian variations, whereas SD1 represents multiple factors influencing skin perfusion (Bottaro et al 2020, Abid et al 2024). In shorter recordings (e.g. 1-hour duration), SD2 likely reflects skin temperature fluctuations due to perfusion changes within that shorter timescale (Abid et al 2024). Since the conventional Poincaré plot does not specify the frequency of oscillations (unlike spectral methods such as Fourier or wavelet analysis), some investigators have used the Extended Poincaré Plot to explore short- and long-term STV across multiple scales (Bottaro et al 2020, Abid et al 2024). The extended version generalises the traditional Poincaré plot by comparing RR intervals separated by a variable lag rather than consecutive ones (Satti et al 2019).
3.6.3. Wavelet transformationTo explore the frequency of physiological oscillations, wavelet transformation can be used to decompose a signal into components of different frequencies (scales). This approach was applied in one of the selected studies (Papaioannou et al 2012) to identify oscillations reflecting neurogenic and metabolic contributions to STV.
3.6.4. Fractal-like exponentsMany physiological signals exhibit scale-free or fractal-like dynamics. DFA is a classic method used to quantify such fractal-like correlations in physiological time-series. It examines the relationship between the logarithm of the scale and the logarithm of variability after removing trends at different scales (Peng et al 1995). A linear relationship between log(scale) and log(variability) indicates a fractal-like structure, and the slope of this line (α) represents the fractal exponent. This method of STV analysis was used in studies by Varela et al (2006
Comments (0)