An EHR-based framework for modeling growth curves and constructing growth centile charts for genetic disorders

Figure 4 summarizes our EHR-based framework for modeling growth curves and centiles, comprising six analytic steps designed to be broadly applicable across genetic conditions and adaptable to different EHR systems.

Fig. 4: Overview of proposed EHR-based framework for modeling growth curves and growth centiles.Fig. 4: Overview of proposed EHR-based framework for modeling growth curves and growth centiles.

EHR electronic health records, GAMLSS Generalized Additive Models for Location, Scale, and Shape; growthcleanr R package, SITAR SuperImposition by Translation and Rotation, BIC Bayesian Information Criterion, OMIM Online Mendelian Inheritance in Man knowledgebase.

Step 1. Cohort definition and data curation

This cohort study was based on patients seen at Vanderbilt University Medical Center (VUMC) between January 1, 2002, and December 31, 2023. Our cohort included patients between the ages of 2 and 20 with a genetically confirmed diagnosis and a minimum of three encounters at least 90 days apart. This study was approved by the Vanderbilt University Medical Center Institutional Review Board (IRB #171011) as exempt research. All procedures were performed in accordance with the ethical standards of the institutional research committee and with the Declaration of Helsinki. The requirement for informed consent was waived by the Institutional Review Board due to the retrospective nature of the study and secondary use of de-identified data. This study followed the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) reporting guidelines.

Our study used patient data derived from the Research Derivative, an image of VUMC’s EHR38. Within this dataset, patients with genetic disorders were identified in the clinical genetic database, a comprehensive database of clinical genetic testing data results developed at VUMC39, with genetically confirmed diagnoses annotated with identifiers from OMIM6 and/or Orphanet7. To define the set of disorders for analysis, we excluded diagnoses that were too rare (i.e., defined as having fewer than 10 male or 10 female cases), cancer-predisposition syndromes primarily associated with adult malignancy risk, genetic risk factors with low penetrance that do not produce characteristic phenotypes, non-syndromic conditions that affect only a single organ system, and late-onset disorders that typically do not manifest until adulthood. After applying these criteria, we identified 15 genetic disorders for sex-specific growth modeling (Table 1).

To compile a list of known growth abnormalities associated with genetic disorders, we used clinical descriptions from OMIM and Orphanet. These descriptions have been mapped to the Human Phenotype Ontology (HPO), a standardized vocabulary designed to describe phenotypes observed in rare and genetic disorders40. We identified all descendants of the HPO terms for disorders of growth (HP:0001507) and puberty (HP:0100000) and linked them to disease IDs in OMIM and Orphanet. To assess the overall relevance of growth abnormalities in rare and genetic disorders, we analyzed the proportion of OMIM and Orphanet diseases that included these terms.

We categorized height-related HPO terms into three domains: size (shorter versus taller), timing (earlier versus later), and intensity (lower versus higher). For example, “Short stature” (HP:0004322) was classified as shorter size, “Delayed puberty” (HP:0000823) as later timing, and “Postnatal growth retardation” (HP:0008897) as lower intensity. When disease descriptions were available in both OMIM and Orphanet, we merged the HPO terms from the two knowledgebases and analyzed their overlap.

Step 2. Data cleaning and preprocessing

Demographic information, including sex and date of birth, was retrieved from the person table in the Observational Medical Outcomes Partnership (OMOP) common data model, which standardizes EHRs for research use. Height measurements ascertained between 2002/01/01 and 2024/03/31 were retrieved from the measurement table in OMOP. We used growthcleanr41, a validated data cleaner for anthropometric measurements in pediatric EHR data, to remove implausible height measurements. Briefly, these included measurements with extreme values, carried-forward entries, or duplicates. We further applied a model-based cleaning procedure using SITAR25, a nonlinear mixed effects model for longitudinal growth data, and excluded observations with standardized residuals exceeding ±49. Because some monogenic conditions are associated with extreme but valid growth patterns, we manually reviewed excluded measurements and identified three patients whose data were removed due to extreme height deviations: Marfan syndrome (n = 2), and Prader-Willi syndrome (n = 1). These individuals were added back into the final dataset for analysis.

Step 3. Generate growth centile charts using GAMLSS

We estimated sex-specific, condition-specific height centiles from ages 2 to 20 using Generalized Additive Models for Location, Scale, and Shape (GAMLSS)24, the standard framework for constructing growth centiles and an extension of the LMS method42. Let \(Y\) denote height\(.\) Height measurements were treated as independent cross-sectional observations. To mitigate overrepresentation of individuals with dense longitudinal follow-up, we retained one observation per individual within each six-month age interval prior to GAMLSS fitting. We assumed that \(Y\sim _\left(y| \mu ,\sigma ,\nu ,\tau \right)\), where \(_\) is a parametric distribution, and \(\mu ,\sigma ,\nu ,\) and \(\tau\) the parameters for location, scale, skewness, and kurtosis, respectively. We modeled each parameter as a smooth function of age using penalized B-splines. We fitted multiple candidate GAMLSS models, including four distributions commonly used in growth modeling (normal, Box-Cox Cole and Green42, Box-Cox power exponential, and Box-Cox \(t\)) and alterative age transformations (log or square root). The smoothness of the age functions was tuned by cross-validation, and the final optimal model was selected by minimizing the Bayesian Information Criterion (BIC)43. Model fit was assessed using worm plots and quantile-quantile plots of residuals. As an additional descriptive check of fit and clinical plausibility, we visually compared the fitted centiles (5th, 10th, 25th, 50th, 75th, 90th, and 95th percentiles) with the observed height data across age. Analyses were performed using the gamlss package in R (v.4.4.2)44,45.

Step 4. Model growth curves using SITAR

We estimated sex-specific, condition-specific height curves from ages 2 to 20 using the SuperImposition by Translation And Rotation (SITAR)25 model, a nonlinear mixed effects model for longitudinal growth data. SITAR models the mean growth trajectory as a natural cubic regression B-spline, in addition to a set of population-level characteristics (fixed effects) and patient-specific effects (random effects). These effects correspond to three clinically interpretable aspects of growth: size, timing, and intensity10,11. For each genetic condition, we fitted SITAR models separately for boys and girls using data from both affected and unaffected individuals. To compare affected and unaffected cohorts, we included cohort status as fixed effects on the size, timing, and intensity parameters, such that the corresponding coefficients directly estimate the mean differences between affected and unaffected cohorts. The SITAR model is shown in Eq. (1):

$$_}=_+_}}+}}_+H\left(\frac_\right)-_-_}}-}}_}_-_}}-}}_\right)}\right)+}}_},$$

(1)

where \(_}\) denotes the height measurement of individual \(i\) at age \(_.\) The fixed effects for size, timing, and intensity are denoted by \(_,_,\) and \(_\), respectively, the random effects by \(}}_,}}_,}}_\) and the cohort status fixed effects by \(_}},_}},\) \(_}}.\) The function \(g\left(\cdot \right)\) is the age transformation, and \(H\left(\cdot \right)\) is the mean growth trajectory modeled using a natural cubic B-spline. Residual errors \(}}_}\) were assumed to be normally distributed, and the random effects were assumed to follow a multivariate normal distribution. We fitted multiple candidate SITAR models varying the spline degrees of freedom from four to nine and allowing alternative age transformations (log or square root). The final optimal model was selected by minimizing BIC43. Model fit was assessed using residual diagnostic plots and quantile-quantile plots of residuals and random effects. Analyses were performed using the sitar and nlme packages in R (v.4.4.2)44.

Step 5. Compare size, timing, and intensity between affected and unaffected cohorts

We quantified differences in growth between affected and unaffected individuals using the SITAR fixed-effect estimates for cohort status. For each genetic condition and sex, these coefficients represent the mean differences in size, timing (age at peak height velocity), and intensity (peak height velocity) between cohorts. Point estimates and corresponding 95% Wald confidence intervals (CI) were obtained from the fitted models35.

Step 6. Validate against annotations on growth abnormality from knowledge bases

We compared SITAR-derived differences in size, timing, and intensity between affected and unaffected individuals to growth abnormality annotations in OMIM and Orphanet. The unit of analysis was the condition–sex–domain combination, where domain corresponded to size, timing, or intensity. We defined annotation-positive controls as condition-sex-domain combinations for which a growth abnormality has already been documented in OMIM or Orphanet. These serve as “known positives” that the SITAR-fitted models should ideally recover. On the other hand, an annotation-negative comparison refers to a condition-sex-domain combination for which no growth abnormality has been documented in these resources. To illustrate with an example, for Marfan syndrome, OMIM and Orphanet have the annotation “tall stature,” which corresponds to an abnormality in size. Therefore, Marfan syndrome-male-size is treated as an annotation-positive control, and we assess whether the SITAR model’s coefficient estimate for cohort status (unaffected vs. affected) for size is directionally correct and statistically significant (i.e., concordant with OMIM/Orphanet). Conversely, Marfan syndrome-male-timing has no existing annotations in OMIM or Orphanet and is therefore treated as an annotation-negative comparison. If the SITAR model’s timing parameter estimate for the difference between affected vs. unaffected cohorts are not statistically significant, we classify this as concordant with OMIM/Orphanet; otherwise, we classify this as a new growth pattern not represented in OMIM or Orphanet. All statistical analyses were performed using R (v.4.4.2)44. The annotated- positive and negative datasets are provided in the Supplementary Information.

Comments (0)

No login
gif