A novel glycolipid composite index predicting cardiovascular disease in Chinese adults with abnormal glucose metabolism: a nationwide cohort study

Abstract

Background:

Cardiovascular disease (CVD) is the leading cause of mortality among individuals with abnormal glucose metabolism. Existing insulin resistance (IR) surrogate indexes show limited predictive capacity in Chinese populations and fail to capture comprehensive glycolipid metabolic dysregulation. We developed and validated TyG-GLM6, a composite index integrating six metabolic parameters (fasting glucose, triglycerides, HDL-C, LDL-C, Age, and BMI), and compared its predictive performance against nine conventional IR indexes.

Methods:

This prospective cohort study analyzed 3,684 participants aged ≥45 years with abnormal glucose metabolism from the China Health and Retirement Longitudinal Study (2011–2020). Associations between TyG-GLM6 and incident CVD were evaluated using multivariate logistic regression and restricted cubic splines. Seven machine learning algorithms were implemented, with performance assessed via ROC curves and SHAP analysis. External validation was conducted in 2,105 participants from a tertiary hospital.

Results:

During 9-year follow-up, 824 (22.4%) participants developed CVD. After full adjustment including biochemical markers, TyG-GLM6 was the only index retaining independent predictive significance (OR: 1.04, 95% CI: 1.01–1.08, P = 0.028), while eGDR, TyG-WC, and CVAI were attenuated to non-significance. TyG-GLM6 exhibited a linear dose-response relationship with CVD risk (P for nonlinear = 0.768, P for overall < 0.001) and consistent performance across sex and age subgroups. Logistic regression achieved optimal performance (AUC: 0.587), with TyG-GLM6 among top predictors. External validation confirmed independent prediction (adjusted OR: 2.02, 95% CI: 1.84–2.20, P < 0.001).

Conclusions:

TyG-GLM6 demonstrates superior independent predictive value for CVD in Chinese adults with abnormal glucose metabolism, outperforming conventional IR indexes in fully adjusted models. Its linear dose-response relationship, demographic robustness, and external validation support its utility for early risk stratification and personalized prevention strategies. Validation in ethnically diverse populations is warranted.

Introduction

Cardiovascular disease (CVD) remains the leading cause of morbidity and mortality worldwide, particularly among elderly populations, where it is characterized by high incidence, disability, and mortality rates, posing persistent challenges to public health systems and socioeconomic development (1, 2). In China, accelerated population aging coupled with lifestyle transitions has driven a rising prevalence of glucose metabolism disorders—including prediabetes and type 2 diabetes—among middle-aged and elderly individuals. Substantial evidence indicates that these populations face significantly elevated CVD risk (36). The underlying mechanisms are closely linked to multiple metabolic dysregulations, including insulin resistance (IR), dyslipidemia, and chronic low-grade inflammation, underscoring abnormal glucose metabolism as a critical CVD risk factor (7). Consequently, developing risk assessment tools that combine strong predictive performance with clinical practicality is essential for enabling early identification and targeted intervention in this high-risk population.

Currently, several IR surrogate indexes based on routine clinical parameters—such as the Triglyceride-Glucose Index (TyG) and estimated Glucose Disposal Rate (eGDR)—are utilized for CVD risk prediction (810). These indexes leverage readily available data from physical examinations, including lipid profiles and glucose levels, offering non-invasive, simple, and cost-effective advantages. Multiple studies have demonstrated their associations with CVD event risk (11, 12). However, existing indexes face notable limitations when applied to middle-aged and elderly Chinese populations with abnormal glucose metabolism. First, their predictive efficacy in this specific cohort remains inadequately validated, raising concerns regarding potential variations related to ethnicity and metabolic background. Second, most indexes focus on isolated metabolic pathways, failing to comprehensively capture the complex, synergistic dysregulation of glucose and lipid metabolism characteristic of this population, thereby limiting their predictive accuracy and clinical utility.

To address these research gaps, we developed a novel glycolipid composite index, TyG-GLM6 (Triglyceride-Glucose–Glycolipid Metabolic 6-index). Building upon the traditional TyG index, TyG-GLM6 integrates six key metabolic parameters, including fasting glucose, triglycerides, high-density lipoprotein cholesterol (HDL-C), low-density lipoprotein cholesterol (LDL-C), and body mass index, to provide a more comprehensive assessment of IR severity and overall glycolipid metabolic dysfunction (13). Using data from a large-scale nationwide prospective cohort, this study systematically evaluated the predictive capacity of TyG-GLM6 for CVD risk in Chinese middle-aged and elderly individuals with abnormal glucose metabolism and compared its discriminative performance against nine conventional IR surrogate indexes. Furthermore, by integrating machine learning approaches, we explored the predictive potential of TyG-GLM6 in complex multidimensional data contexts. This study aims to provide robust scientific evidence for early risk stratification, personalized interventions, and optimized cardiovascular health management strategies in this high-risk population.

MethodsStudy design and population

The China Health and Retirement Longitudinal Study (CHARLS) is an ongoing, nationally representative longitudinal survey designed to assess the social, economic, and health status of Chinese adults aged 45 years and older. The CHARLS cohort was established using a multistage probability sampling strategy, with participants selected from 150 counties (districts) and 450 villages (communities) across 28 provinces nationwide (14). The baseline CHARLS survey was conducted in 2011, with subsequent follow-up waves in 2013, 2015, 2018, and 2020. The CHARLS protocol was approved by the Biomedical Ethics Review Committee of Peking University (IRB00001052-11015), and all participants provided written informed consent prior to enrollment. The study adheres to the principles of the Declaration of Helsinki. Detailed information about CHARLS is publicly available at http://charls.pku.edu.cn/en.

This prospective cohort study utilized data from all five CHARLS survey waves (2011–2020), with 2011 serving as baseline. Follow-up interviews were conducted biennially through standardized face-to-face visits by trained interviewers using computer-assisted technology to ensure data quality and consistency. The baseline survey enrolled 17,705 participants. Participants were sequentially excluded based on the following criteria: (1) age <45 years at baseline; (2) prevalent CVD (heart disease or stroke) or cancer at baseline; (3) absence of abnormal glucose metabolism (prediabetes or type 2 diabetes) at baseline; (4) missing data for TyG-GLM6 or any other IR surrogate index at baseline; (5) incomplete anthropometric, lifestyle, sociodemographic, or biochemical data at baseline; and (6) missing CVD outcome data during follow-up. After applying these criteria, 3,684 participants were included in the final analysis. The participant selection process is illustrated in Figure 1.

Flowchart showing participant selection in CHARLS 2011 with N equals seventeen thousand seven hundred five at baseline, exclusions for age, health status, missing data, and final analytic sample of N equals three thousand six hundred eighty four.

Flowchart of participant selection.

Data collection and measurement

At baseline, trained interviewers collected data using standardized questionnaires, including: (1) demographic characteristics (sex, age, educational attainment, and marital status); (2) anthropometric measurements [systolic blood pressure (SBP), diastolic blood pressure (DBP), height, weight, and waist circumference (WC)]; (3) lifestyle factors (smoking and alcohol consumption status); (4) medical history (hypertension, diabetes, and heart disease); and (5) biochemical parameters [triglycerides (TG), total cholesterol (TC), high-density lipoprotein cholesterol (HDL-C), low-density lipoprotein cholesterol (LDL-C), serum creatinine (Scr), fasting plasma glucose (FPG), glycated hemoglobin A1c (HbA1c), serum uric acid (SUA), blood urea nitrogen (BUN), and C-reactive protein (CRP)].

Educational attainment was dichotomized as below high school vs. high school or above. Smoking status was categorized as current smoker or non-smoker. Alcohol consumption was classified as current drinker or non-drinker. Blood pressure was measured three times after participants rested in a seated position for 5 min, and the average of the three measurements was used for analysis. Venous blood samples were collected after an overnight fast of at least 8 h for all laboratory assessments.

Definition of abnormal glucose metabolism

Prediabetes was defined as FPG 5.6–6.9 mmol/L (100–125 mg/dL) or HbA1c 5.7%–6.4% in the absence of a diabetes diagnosis. Type 2 diabetes was defined as self-reported physician diagnosis, use of glucose-lowering medication, FPG ≥7.0 mmol/L (126 mg/dL), or HbA1c ≥ 6.5%. Abnormal glucose metabolism was defined as the presence of either prediabetes or type 2 diabetes (15). Hypertension was defined as self-reported physician diagnosis, current use of antihypertensive medication, or measured SBP ≥140 mmHg or DBP ≥90 mmHg (16).

Definition of TyG-GLM6 and other IR surrogate indexes

IR was assessed using several validated surrogate indexes derived from routine clinical parameters. The primary index evaluated was TyG-GLM6, which integrates six key metabolic parameters: fasting glucose, triglycerides, LDL-C, HDL-C, waist circumference, and body mass index (BMI). For comparative analysis, nine additional IR surrogate or glycolipid metabolism indexes were included: GLM6, eGDR, TyG, TyG-WC, TyG-BMI, TyG-WHtR, TG/HDL-C, METS-IR, and AIP. All indexes were calculated using formulas from previously published studies, with detailed computational methods provided in Supplementary Table S1.

Outcome ascertainment

The primary outcome was incident CVD, defined as new-onset heart disease or stroke during follow-up. Participants self-reported physician-confirmed diagnoses of CVD at each follow-up wave (2013–2020), consistent with established CHARLS methodology. Incident CVD events were defined as new cases occurring between baseline (2011) and the most recent follow-up (2020) among participants free of CVD at baseline. CVD status was assessed using standardized questions at each wave: “Have you been diagnosed by a physician with heart disease/stroke?” or “Are you currently receiving treatment (traditional Chinese medicine, Western medicine, physical therapy, acupuncture, or occupational therapy) for heart disease/stroke?” The CHARLS research team implemented rigorous quality control procedures to ensure data accuracy and reliability (14).

Missing data handling

Participants with missing data for IR surrogate indexes, CVD outcomes during follow-up, or baseline covariates (sociodemographic and health-related variables) were excluded from the analysis. To assess potential selection bias, baseline characteristics were compared between excluded and included participants.

Machine learning model development and SHAP interpretability analysis

Seven machine learning algorithms were evaluated: Logistic Regression (LR), Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), XGBoost, LightGBM, and Deep Neural Network (DNN). LR estimates binary event probabilities through linear combinations and a sigmoid function; DT partitions data through recursive splitting rules, offering high interpretability but susceptibility to overfitting; RF aggregates multiple decision trees to reduce variance and enhance robustness against noise and outliers; SVM performs classification using maximum margin principles and kernel functions; XGBoost and LightGBM are gradient boosting frameworks that iteratively optimize weak learners—XGBoost excels in handling bias-variance tradeoffs, while LightGBM improves computational efficiency through histogram-based optimization; DNN automatically extracts features through multilayer nonlinear transformations.

All continuous features were standardized to eliminate scale effects. Model hyperparameters were optimized using five-fold cross-validation combined with grid search and manual fine-tuning, with area under the curve (AUC) as the optimization metric. To elucidate model decision-making mechanisms, we employed the SHapley Additive exPlanations (SHAP) framework based on cooperative game theory. SHAP quantifies the marginal contribution of each feature to model predictions by assigning SHAP values, thereby enhancing model interpretability.

Model performance evaluation

Model performance was assessed using receiver operating characteristic (ROC) curves, AUC, sensitivity, specificity, accuracy, and F1 score. Pairwise comparisons of AUC values were conducted using the DeLong test. Decision curve analysis (DCA) and calibration curves were employed to evaluate clinical utility and calibration of the final model. Model calibration was assessed using the Hosmer–Lemeshow test, with P > 0.05 indicating acceptable agreement between predicted and observed event rates.

Statistical analysis

All statistical analyses were performed using R software (version 4.4.3), and machine learning models were developed using the Python scikit-learn library. Two-tailed P < 0.05 was considered statistically significant. Continuous variables are presented as mean ± standard deviation (SD) or median [interquartile range (IQR)] according to their distribution, while categorical variables are presented as frequency (percentage). Between-group comparisons for normally distributed continuous variables were performed using independent t-tests or one-way analysis of variance (ANOVA); non-normally distributed continuous variables were compared using Mann–Whitney U tests. Categorical variables were compared using chi-square tests. Logistic regression models were used to assess associations between TyG-GLM6 and other glycolipid metabolism indexes with incident CVD risk. Cox proportional hazards regression models were conducted as complementary survival analyses. Time-to-event was calculated from baseline (2011) to the first occurrence of CVD or the last follow-up. Schoenfeld residuals and graphical tests were used to examine the proportional hazards assumption. Hazard ratios (HRs) and 95% confidence intervals (CIs) were estimated for each metabolic index.

ResultsBaseline characteristics of study participants

This study included 3,684 participants with abnormal glucose metabolism. During the median follow-up of 9 years (2011–2020), 824 (22.4%) participants developed CVD, while 2,860 (77.6%) remained free of CVD. Baseline characteristics stratified by incident CVD status are presented in Table 1. Compared with participants who did not develop CVD, those with incident CVD were significantly older and had higher values for weight, BMI, WC, BUN, fasting glucose, TG, HbA1c, and all IR surrogate indexes (GLM6, TyG, TyG-GLM6, TyG-WC, TyG-BMI, TyG-WHtR, TG/HDL-C, METS-IR, AIP, CVAI, NHHR, and UHR), as well as lower eGDR and HDL-C levels (all P < 0.05). The proportion of males, current alcohol drinkers, and individuals with hypertension was also significantly higher in the CVD group (all P < 0.05). No significant differences were observed between groups in height, serum creatinine, total cholesterol, LDL-C, CRP, uric acid, educational attainment, or smoking status (all P > 0.05).

VariablesTotal (n = 3,684)Non-CCVD (n = 2,860)New onset CCVD (n = 824)PAge, (years)59.60 ± 9.1459.39 ± 9.2460.33 ± 8.750.010Gender, n(%)0.004Female1,935 (52.52)1,466 (51.26)469 (56.92)Male1,749 (47.48)1,394 (48.74)355 (43.08)Education, n(%)0.600High school and below3,637 (98.72)2,825 (98.78)812 (98.54)High school above47 (1.28)35 (1.22)12 (1.46)Smoking, n(%)0.058No2,256 (61.24)1,728 (60.42)528 (64.08)Yes1,428 (38.76)1,132 (39.58)296 (35.92)Drinking, n(%)0.012No2,404 (65.26)1,836 (64.20)568 (68.93)Yes1,280 (34.74)1,024 (35.80)256 (31.07)Hypertension, n(%)<.001No2,531 (68.70)2,006 (70.14)525 (63.71)Yes1,153 (31.30)854 (29.86)299 (36.29)Height, (cm)157.89 ± 8.74157.94 ± 8.61157.72 ± 9.210.517Weight, (cm)59.55 ± 11.5959.00 ± 11.1861.44 ± 12.72<.001BMI23.82 ± 3.9623.59 ± 3.7724.63 ± 4.46<.001WC, (cm)85.45 ± 11.9184.85 ± 11.5887.54 ± 12.78<.001BUN, (mg/dL)15.95 ± 4.6416.04 ± 4.6915.67 ± 4.450.043GLU, (mg/dL)121.20 ± 38.65120.30 ± 37.55124.35 ± 42.120.013CREA, (mg/dL)0.79 ± 0.270.79 ± 0.280.77 ± 0.180.077CHO, (mg/dL)198.93 ± 39.67198.38 ± 39.63200.86 ± 39.770.113TG, (mg/dL)142.64 ± 108.28140.36 ± 109.78150.57 ± 102.550.017HDL-C, (mg/dL)50.78 ± 15.6251.26 ± 15.6149.12 ± 15.55<.001LDL-C, (mg/dL)119.48 ± 37.24118.97 ± 36.92121.24 ± 38.300.123CRP, (mg/L)2.85 ± 7.722.80 ± 7.873.00 ± 7.190.531HBA1c, (%)5.46 ± 0.985.43 ± 0.965.55 ± 1.040.003UA, (mg/dL)4.50 ± 1.264.51 ± 1.274.47 ± 1.250.421GLM65.70 ± 0.405.68 ± 0.405.78 ± 0.41<.001eGDR9.39 ± 2.169.51 ± 2.118.99 ± 2.26<.001TyG8.85 ± 0.668.82 ± 0.658.93 ± 0.69<.001TyG-GLM650.69 ± 7.1750.34 ± 7.0451.91 ± 7.47<.001TyG-WC757.70 ± 131.67750.28 ± 128.48783.46 ± 139.22<.001TyG-BMI211.44 ± 42.13208.77 ± 40.46220.71 ± 46.31<.001TyG-WHtR4.81 ± 0.854.76 ± 0.834.98 ± 0.89<.001TG/HDL-C3.56 ± 4.753.47 ± 4.723.89 ± 4.850.025METS-IR6.90 ± 0.546.88 ± 0.546.99 ± 0.55<.001AIP0.89 ± 0.800.86 ± 0.801.00 ± 0.80<.001CVAI99.29 ± 44.1296.35 ± 43.35109.52 ± 45.26<.001NHHR3.26 ± 1.563.20 ± 1.523.46 ± 1.65<.001UHR0.10 ± 0.050.10 ± 0.050.10 ± 0.060.015

Baseline characteristics of the study participants.

Bold value indicates P < 0.05.

Association of IR surrogate indexes with incident CVD

To examine associations between IR surrogate indexes and incident CVD among participants with abnormal glucose metabolism, univariate and multivariate logistic regression analyses were performed (Table 2). Univariate analysis revealed that GLM6 (OR: 1.88, 95% CI: 1.68–2.11), eGDR (OR: 0.89, 95% CI: 0.86–0.92), TyG (OR: 1.28, 95% CI: 1.17–1.40), TyG-GLM6 (OR: 1.03, 95% CI: 1.02–1.04), TyG-WC (OR: 1.01, 95% CI: 1.01–1.01), TyG-BMI (OR: 1.01, 95% CI: 1.00–1.01), TyG-WHtR (OR: 1.36, 95% CI: 1.23–1.51), TG/HDL-C (OR: 1.02, 95% CI: 1.01–1.03), AIP (OR: 1.23, 95% CI: 1.13–1.34), CVAI (OR: 1.01, 95% CI: 1.01–1.01), NHHR (OR: 1.11, 95% CI: 1.08–1.15), and UHR (OR: 6.17, 95% CI: 4.76–7.99) were all significantly associated with incident CVD (all P < 0.05). In multivariate analysis adjusting for demographic characteristics, lifestyle factors, and anthropometric measurements, eGDR (OR: 0.94, 95% CI: 0.90–0.99), TyG-GLM6 (OR: 1.02, 95% CI: 1.00–1.03), TyG-WC (OR: 0.99, 95% CI: 0.99–1.00), and CVAI (OR: 1.01, 95% CI: 1.00–1.01) remained independently associated with incident CVD (all P < 0.05). These findings indicate that the novel glycolipid composite index TyG-GLM6 is significantly associated with CVD risk in this population.

VariablesUnivariateMultivariateβS.EZPOR (95%CI)βS.EZPOR (95%CI)GLM60.630.106.36<.0011.88 (1.55–2.29)eGDR−0.110.02−6.12<.0010.89 (0.86–0.93)−0.060.02−2.540.0110.94 (0.90–0.99)TyG0.250.064.29<.0011.28 (1.14–1.44)TyG-GLM60.030.015.52<.0011.03 (1.02–1.04)0.020.012.600.0091.02 (1.01–1.03)TyG-WC0.010.006.35<.0011.01 (1.01–1.01)−0.010.00−2.410.0160.99 (0.99–0.99)TyG-BMI0.010.007.00<.0011.01 (1.01–1.01)TyG-WHtR0.310.056.42<.0011.36 (1.24–1.50)TG/HDL0.020.012.210.0271.02 (1.01–1.03)AIP0.200.054.18<.0011.23 (1.11–1.35)CVAI0.010.007.53<.0011.01 (1.01–1.01)0.010.004.28<.0011.01 (1.01–1.01)NHHR0.100.024.15<.0011.11 (1.05–1.16)UHR1.820.762.390.0176.17 (1.39–27.51)

Univariate and multivariate regression analysis of potential indicators.

OR, odds ratio; CI, confidence interval.

Bold value indicates P < 0.05.

Dose–response relationships between IR surrogate indexes and incident CVD

Table 3 presents the associations between the four significant multivariate predictors (TyG-GLM6, eGDR, TyG-WC, and CVAI) and incident CVD across three progressively adjusted models. In the unadjusted model (Model 1), all four indexes were significantly associated with CVD risk (all P < 0.001). After adjustment for demographic characteristics, lifestyle factors, and anthropometric measurements (Model 2), TyG-GLM6 (OR: 1.02, 95% CI: 1.01–1.03, P = 0.007), eGDR (OR: 0.94, 95% CI: 0.90–0.99, P = 0.017), TyG-WC (OR: 1.01, 95% CI: 1.01–1.01, P = 0.024), and CVAI (OR: 1.01, 95% CI: 1.01–1.01, P = 0.027) all retained statistical significance. However, after further adjustment for all biochemical markers (Model 3), only TyG-GLM6 remained independently associated with incident CVD (OR: 1.04, 95% CI: 1.01–1.08, P = 0.028), while the other three indexes were attenuated to non-significance.

VariablesModel1Model2Model3OR (95%CI)POR (95%CI)POR (95%CI)PTyG-GLM61.03 (1.02–1.04)<.0011.02 (1.01–1.03)0.0071.04 (1.01–1.08)0.028eGDR0.89 (0.86–0.93)<.0010.94 (0.90–0.99)0.0170.93 (0.76–1.14)0.493TyG-WC1.01 (1.01–1.01)<.0011.01 (1.01–1.01)0.0241.00 (1.00–1.01)0.217

Comments (0)

No login
gif