Handgrip strength (HGS) is broadly recognized as a reliable and non-invasive biomarker of overall health. It is typically measured isometrically using a hydraulic hand dynamometer, an instrument that shows good reassessment reliability [1, 2]. This measurement involves participants squeezing the dynamometer with maximum effort without any hand or arm movement, thus measuring the isometric grip force. Direct measured HGS offers several advantages, including low cost, ease of administration, and a strong predictive value for various health outcomes [1, 3–6]. Beyond, HGS is related to the gray matter volume (GMV) of the brain, which in turn has been used as a marker for neuropathological changes in neurodegenerative and psychiatric diseases. GMV has also been associated with protective factors such as muscular strength [1]. Lower GMV is associated with lower HGS [7]. On the other hand, stronger HGS is associated with higher GMV in a wide array of brain regions like the ventral striatum, hippocampus, thalamus, pallidum, putamen, brain stem, temporal pole, and parahippocampal gyrus [1]. Through its associations with physical capabilities and with structural brain integrity, HGS offers insights into the neurobiological mechanisms related to muscle strength [1]. This relationship, however, is influenced by genetics, physical fitness, mental health as well as anthropometrics, and demographic characteristics, emphasizing the need for a comprehensive examination of these associations [1].
HGS serves as a marker of overall health because it can be affected through different mechanisms. Lower HGS in older adults is associated with adverse health outcomes, such as increased mortality risk, diminished functional mobility, cognitive impairments, and a range of health issues, including metabolic diseases like diabetes, and neurological conditions including stroke and major depressive disorder (MDD) [8–10]. Moreover, reduced HGS is not only linked to higher disease recovery times [11] but also functions as a valuable biomarker for assessing recovery and prognosis in stroke patients, with evidence linking a decrease in HGS after stroke (post-stroke) compared to before stroke (pre-stroke) levels [12, 13]. These attributes have established HGS as an indicator of muscle strength and general health in clinical settings [1, 8, 11, 14]. While understanding longitudinal changes in HGS in patients can help risk assessment and improve monitoring, to date it remains understudied due to lack of longitudinal data.
Anthropometric factors such as height, body mass index (BMI), and waist-to-hip ratio (WHR) are directly linked to HGS [15–20]. For instance, BMI and height correlate positively with greater strength [17, 19]. Conversely, WHR, which reflects the proportion of abdominal fat relative to hip circumference, often exhibits an inverse relationship with HGS. Greater abdominal fat relative to hip size is frequently associated with lower HGS [18]. Age is another critical factor affecting HGS. Generally, HGS increases during childhood and adolescence, peaks in early adulthood, and then declines with higher age, particularly after 40. This decline is often more pronounced in older adults, especially those over 75, where the rate of decrease accelerates [16]. The aging process results in progressive muscle loss, further impacting HGS [16, 21]. Sex differences also contribute significantly to variations in HGS [22, 23], influencing both baseline strength levels and decline patterns. Males and females exhibit distinct patterns in disease progression and anthropometric measurements, necessitating sex-specific considerations in clinical and research settings [24]. The interpretation of HGS is thus most meaningful when normalized to anthropometric and demographic factors, rather than relying on raw measured values. Relative HGS, which accounts for these variables, may offer a more precise assessment of an individual’s neuromuscular deficit and its relationship with brain structure and diseases, potentially improving its utility as a health indicator [8, 11, 12, 25]. In particular, a refined HGS interpretation can open up new possibilities for early disease detection, monitoring treatment efficacy, guiding recovery processes, and predicting long-term health outcomes across various conditions.
Machine learning (ML) based predictive modeling offers individual-level predictions, establishing ML as a transformative tool in modern clinical practice [20, 26]. ML can be used to predict HGS using anthropometric and demographic features. This predicted HGS captures the variance explained by the features and thus can be used to develop a relative HGS score. In this study, we tested the hypothesis that the difference between true HGS and predicted HGS can serve as a biomarker for muscular strength relating to the structural integrity of the brain and disease-related changes. Using data from the UK Biobank (UKB), we developed sex-specific ML models to predict HGS using anthropometrics and demographic variables in healthy individuals. We applied statistical bias-correction techniques to enhance prediction accuracy and introduced a novel score called
, defined as the difference between the true HGS and bias-free predicted HGS. We then investigated the brain basis
based on correlation with regional GMV followed by the investigation of its sensitivity to longitudinal changes in patient cohorts from two diseases known to influence HGS: stroke and MDD. The key innovation lies in our longitudinal design, capturing HGS scores at two time points: before (pre) and after (post) disease onset. This approach enabled us to examine how
and true HGS change over time in patients compared to healthy controls (HC) groups.
We used data from the UKB database with more than half a million adult participants recruited from a total of 22 assessment centers across the United Kingdom with baseline assessment between 2006–2010. The baseline assessment included a wide range of demographic data, physical measurements, clinical and health-related information, and the completion of a touchscreen questionnaire [27, 28]. A subset of participants was invited back in 2012 and 2013 for the first repeat assessment. During this repeat visit, additional data were collected, although no brain imaging data were collected at either the baseline or the first repeat assessment. Starting in 2014, a subsample was invited to assessment centers for brain imaging, with follow-up imaging assessments beginning in 2019. In this study, participants were categorized into three groups: HC and two patient cohorts: stroke and MDD. Group definitions were established using inclusion and exclusion criteria based on the international classification of diseases, 10th revision (ICD-10) coding system. Only individuals with all anthropometrics, demographics, and HGS values available were included (figure 1).
Figure 1. The ML analysis pipeline used to predict HGS and analyze its associations with neurobiological markers and disease-related impairments in this study. For data preparation anthropometric (BMI, height, and waist-to-hip ratio) and demographic (age) variables were obtained from the UK Biobank database. These variables (predictors) were used to train sex-specific ML models, specifically linear SVM, random forest (RF) and extreme gradient boosting (XGBoost), to predict HGS. The pipeline includes steps for preprocessing, model training, performance evaluation through cross-validation, and the application of statistical bias-correction techniques to enhance prediction accuracy. The true HGS and
(
; see section 2.2.2) scores are then assessed for their correlation with neurobiological markers, such as gray matter volume (GMV), and their effectiveness in distinguishing between HC and patients with stroke and MDD.
Download figure:
Standard image High-resolution image 2.1.1. Healthy controlsThe HC population was obtained by excluding participants with a known history or current diagnosis of mental and behavioral, psychiatric, nervous system, neurological, cardiovascular, cerebrovascular, musculoskeletal system, connective tissue conditions, injuries or poisoning diseases as outlined in the ICD-10 codes (table 1). The HC participants were divided into two groups: 1) participants who did not undergo brain imaging assessments (non-imaging data,
= 201 133) and 2) participants who participated in at least one brain imaging assessment out of the two available imaging assessments (HC with imaging data,
= 32 125). Individuals with non-imaging data were used to train and evaluate ML models. The HC individuals with imaging data were employed both in assessing the association between HGS and GMV as well as for matched HC comparisons with patient cohorts (supplementary figure S1).
Table 1. Definition of the study populations based on ICD-10 criteria for HC, stroke, and MDD.
PopulationsExcluded ICD-10 criteriaIncluded ICD-10 criteriaHealthy controls (HC)Mental and behavioral disorders: F—Diseases of the nervous system: G Cerebrovascular diseases: I60–I69 Diseases of the musculoskeletal system and connective tissue: M Injury, poisoning, and certain other consequences of external causes: S StrokeMental and behavioral disorders: FIschemic: I63Diseases of the nervous system: G00–G14, G20–G26, G30–G32, G35–G37, G54–G59, G60–G61, G70–G73, G80–G83, G91–G99Intracerebral hemorrhage: I61Diseases of the musculoskeletal system and connective tissue: M Injury, poisoning, and certain other consequences of external causes: S Major depressive disorder (MDD)Mental and behavioral disorders: F00–F31, F34–F48, F50–F99Depressive episode: F32Diseases of the nervous system: GRecurrent depressive disorder: F33Cerebrovascular diseases: I60–I69 Diseases of the musculoskeletal system and connective tissue: MInjury, poisoning, and certain other consequences of external causes: S 2.1.2. Patient cohortsFirst, participants with conditions of the musculoskeletal system, connective tissue, or injury were excluded from all patient sample groups (table 1). The outcomes of incident stroke were defined according to the ‘Algorithmically-defined outcomes (ADOs)’ (UKB Resource 460) developed by the UKB team [29]. The algorithm integrated information from UKB’s baseline assessment data collection along with linkage data, including hospital admissions, diagnoses and procedures, death register records, and self-reported medical condition codes reported at the baseline assessment visit. The incident MDD outcome was obtained from the ‘First occurrence of health outcomes defined by 3-character ICD-10 code’ algorithm (UKB Resource 593) [30, 31]. The UKB indicated the first occurrence of a set of diagnostic codes for a wide range of health outcomes across self-report, primary care, hospital inpatient data, and death data, mapped to a 3-digit ICD-10 code. To establish patient cohorts, we included stroke endpoints comprised of ischemic stroke (I63) or intracerebral hemorrhagic stroke (I61), and MDD depressive episode (F32) or recurrent depressive disorder (F33). The onset dates for the diagnoses of two diseases were identified using the first occurrence fields: stroke (Data-Field 42006), and MDD (Data-Fields 130894 and 130896). To ensure diagnostic accuracy, cases based solely on self-reported data were excluded from the analysis. Patients with a history of diseases prior to their baseline assessment visit were excluded to ensure that the analysis focuses on incident cases. Additional exclusion criteria were applied to each patient group, including missing dates of disease onset, missing data, and relevant HGS conditions (see section 2.1.3). After applying exclusion criteria, we identified disease cohorts consisting of patients with longitudinal data who completed assessments at two time-points: (1) an initial assessment visit prior to the onset of the disease (pre time-point), and (2) the first follow-up assessment visit after disease onset (post time-point). The final patient cohorts comprised: 40 males and 16 females in the stroke group (1 female patient hemorrhagic and all others ischemic), and 37 males and 60 females in the MDD group. These cohorts provided longitudinal data for analysis of HGS changes concerning disease onset.
2.1.3. HGS assessmentHGS was measured isometrically using a calibrated Jamar J00105 hydraulic hand dynamometer, which was monitored by a research assistant. During the HGS measurement, participants were told to sit upright in a chair with their forearms resting on armrests pointing forward and their elbows bent and locked at a 90° angle. The maximum HGS value was obtained from each hand while participants were instructed to squeeze the handle as hard as possible for approximately 3 s. Both hands were measured consecutively (Data-Field 46 for the left and Data-Field 47 for the right hand). Participants whose dominant HGS<4 kg or lower than their non-dominant HGS were eliminated from further analysis [1, 32]. Hand dominance was based on self-report. If information on handedness was not available or if the individual reported using both hands (ambidextrous), we based dominance on the highest HGS score obtained from either right or left hand. In this study, our target of interest was combined HGS (which we refer to simply as HGS), calculated as the sum of the grip strength measurements from the right and left hands. While measuring HGS in each hand separately can reveal unilateral deficiencies, assessing combined HGS provides a comprehensive measure of overall hand strength. In clinical settings, assessing the strength of both hands offers a robust measure of overall strength and helps identify unilateral weaknesses or conditions affecting one side of the body that may be overlooked with single-hand testing, especially for conditions like stroke or localized musculoskeletal disorders [33]. This approach is particularly relevant when considering the 10% rule, which states that the dominant hand typically has a 10% greater grip strength than the non-dominant hand, primarily applies to right-handed individuals, who make up more than 90% of both male and female participants in this study. In contrast, for left-handed individuals, grip strength tends to be more balanced between both hands [34]. Therefore, combined HGS offers a more equitable assessment for left-handed individuals, as it eliminates the need for adjustments based on hand dominance. Note that for lateralized motor deficits as encountered in stroke, the combined HGS score will also be reduced.
2.1.4. Demographic and anthropometric assessmentsParticipants’ sex was determined from self-reported information (Data-Field 31). Age was calculated based on the date of the baseline assessment attendance and the participant’s birth date. Anthropometric data were obtained during the physical measures phase of each assessment visit. Height (Data-Field 50) was directly measured, while BMI (Data-Field 21001) was calculated using weight and height data (kg m−2). WHR was determined by dividing waist circumference (Data-Field 48) by hip circumference (Data-Field 49).
2.2. ML analysis2.2.1. Data preparationData from non-imaging HC participants who only attended the baseline assessment was used to train and evaluate ML models. Age and anthropometric (i.e. BMI, height, and WHR) characteristics were considered as predictors in our models. To avoid overfitting and base decisions on the most promising models, we split the HC data into training (90%) and test datasets (10%). The splits were stratified based on binned age (into 5 bins), HGS (into 5 bins), and sex to keep the splits representative of the whole population. After excluding cases with missing data and relevant HGS conditions (see section 2.1.3) from each dataset, the final HC data included 61 816 males (43.32%) and 80 886 females (56.68%) in the training set, and 6938 males (43.62%) and 8969 females (56.38%) in the test set (table 2, figure 1, and supplementary figure S1).
Table 2. Summary of the HC non-imaging population characteristics.
HC non-brain-imaging at baseline assessment visitDatasetTrain datasetTest datasetSexBoth sexFemaleMaleBoth sexFemaleMaleNumber142 70280 88661 81615 90789696938Age, mean (SD)55.46 (8.12)55.24 (7.99)55.75 (8.27)55.54 (8.13)55.28 (8.01)55.87 (8.26)BMI, mean (SD)26.76 (4.42)26.35 (4.74)27.29 (3.91)26.74 (4.35)26.29 (4.63)27.33 (3.89)Height, mean (SD)168.46 (9.21)162.76 (6.27)175.93 (6.79)168.47 (9.24)162.69 (6.24)175.95 (6.8)WHR, mean (SD)0.86 (0.09)0.81 (0.07)0.93 (0.06)0.86 (0.09)0.81 (0.07)0.93 (0.06)Combined HGS, mean (SD)62.50 (21.27)48.65 (11.68)80.62 (16.92)62.56 (21.25)48.69 (11.69)80.48 (16.99)Right dominant hand90.34%91.93%88.26%90.16%91.6%88.3%2.2.2. Model training and performance evaluationWe utilized the non-imaging HC training dataset (61 816 males and 80 886 females) to train sex-specific ML models for HGS prediction using the anthropometrics and age features. To prevent sex bias and given known sex differences in HGS, and anthropometric features, models were trained separately for males and females. We employed linear support vector machine (SVM), random forest (RF) and extreme gradient boosting (XGBoost) regression models. Our motivation for including the nonlinear predictive models besides the linear SVM in our analysis was based on the fact that a non-linear relationship between age and HGS has already been documented [11, 20, 35]. Specifically, HGS generally increases until approximately ages 30–40, after which it begins to decline, and non-linear models are needed to capture such association [35, 36]. Pearson’s correlation coefficient (
), the coefficient of determination (
) and mean absolute error (MAE) were used to compare model performance. To obtain generalization estimates, we performed 10 times repeated 10-fold (10
10-fold) cross-validation (CV) using the Julearn ML library version 0.2.7 (https://juaml.github.io/julearn/) [37], building on top of the scikit-learn library [38]. The hyperparameter
for the linear SVM was calculated using a heuristic as
where
is the number of subjects,
is the feature matrix of the dataset, and
denotes the value of feature
for subject [39]. As alternative non-linear approaches, we trained both RF and XGBoost regression models using the Scikit-Learn (sklearn) Python package version 0.23.2.
In the first step, we performed a scaling study by comparing the prediction performances of the linear SVM, RF and XGBoost models across six levels of sample sizes (10, 20, 40, 60, 80, and 100%) generated by randomly selecting the required number of sampling data from the training dataset to perform the CV procedure. Comparing model performances across different sample sizes help evaluate model stability and reliability by revealing how performance changes with increasing amounts of data. The model with higher performance was selected for further analysis. Feature importance (FI) scores for each model were derived to quantify the influence of individual feature variables on model outputs. For the linear SVM model, we used the coefficient parameter (.coef_) as the FI scores.
In the second step, we validated the models trained using the whole training data (90% of HC), by comparing the true HGS and predicted HGS (
) values on the 10% hold-out test set. This step is crucial for determining how well the trained models perform and identifying any inherent biases or errors in the prediction process. To this end, we calculated the difference between the true HGS and the
(i.e.
). A positive
indicates that the individual is stronger than expected, while a negative value suggests weaker than expected strength. However, assessing an individual’s strength using this difference is highly dependent on the accuracy of the HGS prediction model. Prediction frameworks often encounter bias, characterized by overestimation of low values and underestimation of high values in the target variable. Such a bias can compromise model accuracy and impair the interpretability of predictions on new data [40]. To address this issue and enhance model performance, we implemented a statistical bias-correction technique, adjusting HGS predictions to align more closely with the true HGS distribution. We applied the bias-correction method developed by Beheshti et al for brain age prediction, given its demonstrated effectiveness in reducing variance [41]. The correction is employed by means of a linear regression model between true HGS and the
. We fitted this model using the predictions from out-of-sample validation sets generated through a 10-fold CV on the training set and calculated the slope (α) and intercept (β), which are used to correct the predictions to achieve a bias-free predicted HGS value (
; equation (1)).

Like the prediction models, the bias correction models were trained separately for males and females. Finally, the proposed biomarker
was defined by subtracting the bias-free predicted HGS (
) from true HGS according to equation (2).
2.3. Reassessment reliabilityWe then assessed the agreement between
values by using the reassessment reliability process in the subset of HC test dataset. To evaluate the reliability of the selected sex-specific trained models, we selected participants from the test dataset for which two non-imaging UKB assessment visit sessions were available: (1) baseline assessment visit as the initial measurement, and (2) first repeat assessment visit (after 2–7 years) as the reassessment, resulting in 134 males and 162 females for this analysis (supplementary table S2). The concordance correlation coefficient (CCC) [42] between
from the two assessment visit sessions was calculated. This reassessment reliability analysis was designed to assess potential changes in relative HGS over time, considering their new age and possible health changes. The analysis reflects both the stability of scores and their sensitivity to genuine changes in participant conditions over time, which are crucial factors for interpreting the reliability of outcomes.
We investigated the neurobiological basis of the HGS scores and their association with GMV. For this, we utilized MRI data from the UKB’s first imaging assessment visit [43], acquired using 3T scanners following the protocol and acquisition parameters detailed in Miller et al [44]. The structural preprocessing of these images was conducted using pipelines developed and executed by the UKB [45]. Specifically, we analyzed the extracted parcel-wise GMV features from T1-weighted (T1w) preprocessed images. The initial preprocessing of the MRI data involved retrieving T1w preprocessed images from the UKB, which were then converted into a DataLad dataset for provenance tracking [46]. Subsequently, voxel-based morphometry was computed using the computational anatomy toolbox version 12.7, resulting in images normalized to the MNI152 space with a 1.5 mm isotropic resolution [47]. The parcel-wise GMV was extracted as the winsorized mean (with limits set at 10%) of the voxel-wise values per parcel, combining three different brain atlases: the Schaefer et al cortical atlas (1000-parcel) [48], the S4 3T version of Tian et al Melbourne subcortical atlas (54-parcel) [49], and the cerebellum SUIT Diedrichsen et al atlas (34-parcel) [50]. The result was a feature vector containing the percolated GMV of 1088 brain regions of interest for each participant. To accommodate for individual differences in total intracranial volume, we linearly regressed it out from each brain region. For this analysis, we used a subset of HC participants who completed the initial imaging visit with 11 077 males and 12 849 females (supplementary table S1). The final samples were reduced to 7726 males and 9292 females based on the availability of all 1088 GMV features.
2.4.2. Correlation analysisWe investigated the correlation between the GMV of 1088 regions separately for true HGS and
. The
values were obtained using the best-performing separate linear SVM models for males and females (see section 2.2.2). Pearson’s
and corresponding
-values were computed. To focus on robust associations, we applied a correlation threshold of
> 0.1 (the absolute values of correlation coefficients exceeding 0.1) together with correcting
-values using the Bonferroni correction with significance determined at
0.05. Separate analyses were conducted for male and female participants to identify regions significantly associated with each HGS score. The correlation coefficients for these significant regions were visualized on brain maps for each sex individually. Additionally, the intersection of the regions showing significance for both sexes were visualized using averaged correlation coefficients.
We identified patient cohorts with longitudinal data from pre (prior to the disease onset) and post (follow-up the disease onset) time-points. The
score could provide insights into disease-specific variation in HGS between pre and post time-points in longitudinal cases. To investigate this, patients with diseases were compared with matched HC samples using a 1:10 (case: control) ratio, ensuring robust comparative analyses. The matching process entailed selecting HC individuals whose assessment visit sessions at both pre and post time points coincided with those of the patients, thereby maintaining temporal consistency.
The final patient cohorts consisted of 40 males and 16 females with stroke, and 37 males and 60 females with MDD (see section 2.1.2). Matched HCs were selected from the pool of HC participants who had undergone brain imaging assessment visits and were not used in model training (15 516 males and 16 609 females). Further refinement of subsamples involved excluding data with missing values and additional HGS conditions (see section 2.1.3). The final number of HC participants included in the matching process for each assessment was as follows: baseline (11 918 males and 13 714 females), first repeat (1884 males and 2013 females), imaging visit (11 077 males and 12 849 females), and first repeat imaging (1153 males and 1359 females). These refined HC cohorts enabled comprehensive comparisons with patient cohorts across multiple assessment time points. The HC subsample used in the matched control-case study, categorized by assessment visit, is detailed in supplementary table S1. Propensity score matching was applied using a 1:10 nearest-neighbor approach to select HC samples for each patient within each disease cohort [51]. This method involved matching each patient with 10 HC participants who had similar propensity scores, taking into account age, anthropometric features at the pre time-point, and the time interval between pre- and post- assessment visits (days). Finally, post time-point data were identified for each HC participant to maintain subject consistency between pre and post time-points. This approach ensured the availability of longitudinal data in the matched HC group, enabling a robust temporal comparison with the patient cohorts. The matching process was conducted without replacement, ensuring that each HC individual could be selected as a match only once per patient group (an overlap was allowed of HC for the different patient samples, i.e. overlaps of HC for stroke and HC for MDD). The characteristics of the matched HC samples and patients summarized in table 3.
Table 3. Characteristics of matched healthy controls (HC) and patients.
SexMaleFemaleGroupPatientMatched HCPatientMatched HCTime-pointPrePostPrePostPrePostPrePostStrokeNumber40404004001616160160Age, mean (SD)59.65 (6.79)68.91 (7.95)59.31 (6.86)68.24 (7.34)61.72 (5.08)71.21 (6.91)61.29 (5.84)71.01 (7.34)BMI, mean (SD)26.7 (3.23)25.85 (3.32)26.68 (3.79)26.53 (4.02)26.37 (5.26)26.7 (5.36)26.68 (4.71)26.36 (4.79)Height, mean (SD)177.22 (7.23)176.44 (7.19)177.34 (6.35)176.71 (6.51)162.39 (7.43)161.3 (7.02)162.32 (5.89)161.34 (6.05)WHR, mean (SD)0.94 (0.06)0.94 (0.06)0.94 (0.06)0.94 (0.06)0.82 (0.1)0.84 (0.1)0.82 (0.07)0.84 (0.07)C
Comments (0)