Objective:
Congenital Hypogonadotropic Hypogonadism (CHH) is a rare disease with an extremely low incidence, and the outcomes of fertility induction therapy in CHH patients exhibit significant interindividual heterogeneity. Given the context of scarce samples and heterogeneous phenotypes, conventional statistical methods struggle to integrate multidimensional data and quantify the contribution of influencing factors. In contrast, machine learning (ML) techniques offer unique advantages in integrating high-dimensional complex medical data and uncovering hidden relationships. To date, the application of ML for predicting treatment outcomes in CHH remains unexplored. Therefore, this study aims to, for the first time, utilize an ML algorithm to construct and validate a predictive model based on a limited clinical cohort and thereby provide a basis for the individualized treatment of CHH.
Methods:
In this single-center retrospective cohort study, 65 adolescent male CHH patients undergoing fertility induction therapy were enrolled and categorized into success (nocturnal emission, n = 55) and failure (non-ejaculation, n = 10) groups based on treatment outcomes. Fifteen pre-treatment baseline indicators across four categories were collected. A random forest model was constructed, employing the Synthetic Minority Over-sampling Technique (SMOTE) and 5-fold stratified cross-validation to mitigate class imbalance and overfitting. Key predictors were identified via feature importance ranking, and decision thresholds were optimized using ROC curves. The model’s performance was comprehensively evaluated and compared against other ML methods.
Results:
The random forest model demonstrated excellent and stable predictive performance: Accuracy 0.84 ± 0.12, Precision 0.82 ± 0.13, Recall 0.89 ± 0.08, F1 score 0.85 ± 0.10, and AUC 0.95 ± 0.04. Feature importance analysis identified the top five predictors: cryptorchidism (the strongest predictor), pre-treatment penile diameter, penile length, follicle-stimulating hormone (FSH) level, and anti-Müllerian hormone (AMH) level. Comparative analysis confirmed the superior comprehensive performance of the random forest model.
Conclusion:
This study successfully developed a robust machine learning model for predicting CHH treatment outcomes. It not only validates the methodological utility of ML in small-sample rare disease research but also elucidates the physiological basis of treatment response through interpretable feature clusters. The model provides clinicians with a quantifiable tool for risk stratification and paves the way for personalized therapeutic decision-making in CHH.
1 IntroductionCongenital Hypogonadotropic Hypogonadism (CHH) is a rare genetic disorder characterized by defects in the Hypothalamic-Pituitary-Gonadal (HPG) axis, with an incidence rate of approximately 1/10,000 to 1/84,000, showing significantly higher prevalence in males than in females (male-to-female ratio≈3–5:1) (Young et al., 2019). The core pathological mechanism involves impaired migration or dysfunction of Gonadotropin-Releasing Hormone (GnRH) neurons, resulting in insufficient gonadotropin secretion, which subsequently causes arrested gonadal development, absence of secondary sexual characteristics, and impaired fertility (Boehm et al., 2015). Currently, the mainstream clinical fertility induction therapy regimens include GnRH pulse therapy and combined hCG/hMG replacement therapy, which have been demonstrated to effectively initiate pubertal development and restore gonadal function in partial patients, achieving final spermatogenesis rates ranging from 64% to 95% across different studies (Young et al., 2019).
Existing research on factors influencing treatment outcomes in male CHH patients has primarily been confined to traditional statistical univariate or multivariate regression analyses. Reported associations include history of cryptorchidism, baseline testicular volume, baseline Inhibin B (INH-b) and anti-Müllerian hormone (AMH) levels, as well as the presence of positive genetic mutations (Rohayem et al., 2017; Cho et al., 2025; Pitteloud et al., 2002; Cangiano et al., 2021). However, such approaches exhibit significant limitations: firstly, they struggle to integrate multi-dimensional high-dimensional data encompassing clinical examinations, serological markers, and genetic data, often overlooking synergistic effects among indicators; secondly, they fail to provide a precise quantitative ranking of contributing factors’ impact magnitudes, making it difficult to meet individualized clinical management needs; thirdly, research subjects were predominantly adults or mixed adult-adolescent cohorts, with scarce studies focusing exclusively on adolescent male patients. The limited sample size constrains the application of statistical methodologies. In recent years, machine learning techniques have rapidly expanded within the medical domain. Their robust capabilities in non-linear fitting and high-dimensional feature mining have demonstrated significant advantages in areas such as diabetic complication prediction and tumor prognosis assessment (Gulshan et al., 2016; Esteva et al., 2017), offering a novel technical pathway for investigating treatment outcomes in rare diseases like CHH.
Artificial Intelligence (AI) has demonstrated translational potential in clinical medical research, with its core value manifested in the integration of high-dimensional heterogeneous data and enhancement of predictive capabilities. In the field of medical imaging, deep learning models can efficiently identify pathological lesions in X-rays and histopathological slides, enabling automated preliminary screening and quantitative analysis (Gulshan et al., 2016; Esteva et al., 2017; Hong et al., 2023; Esteva et al., 2019; Perez et al., 2019). Concurrently, machine learning can be leveraged to integrate multidimensional clinical data (medical records, genomics, vital signs) to construct disease risk prediction models, thereby facilitating early diagnosis and prognostic assessment (Cai et al., 2024; Yamamoto et al., 2022; Lee et al., 2023; Inoue, 2024; Lisik et al., 2025; AlShehhi et al., 2025; Guo et al., 2025). These technologies not only alleviate the burden of repetitive workloads for physicians but also serve as crucial auxiliary tools for enhancing diagnostic consistency and enabling personalized medicine by identifying patterns imperceptible to the human eye. However, in the field of rare diseases, particularly syndromes with exceptionally low incidence rates such as CHH, limited sample size and phenotypic heterogeneity constitute critical bottlenecks. Gonadal disorder follow-up is simultaneously impacted by privacy concerns, presenting considerable challenges in clinical data acquisition. In AI applications, limited by the number of complete follow-up cases, deep learning is highly susceptible to falling into overfitting, which consequently weakens the model’s extrapolation capability. In contrast, classical machine learning algorithms represented by random forest demonstrate superior bias-variance trade-offs in extremely small sample scenarios through their ‘shallow structure + ensemble strategy’: their splitting criterion is insensitive to feature scale and can naturally handle mixed-type variables; By introducing perturbations through bootstrap sampling and feature random subspace, they effectively reduce variance while preserving clinically interpretable pathways (Wang et al., 2022; Chen et al., 2023; Johnson et al., 2022). Therefore, in clinical research on rare diseases, prioritizing decision tree algorithms achieves dual objectives of robust prediction and knowledge discovery under data-constrained conditions (Banerjee et al., 2023). Currently, the application of ML learning in the CHH field remains at a nascent stage, and no relevant reports have been published. Building upon this, our study leveraged a single-center cohort with a relatively large sample size of CHH patients, integrating 15 baseline indicators across four categories to construct multiple ML predictive models. This approach aims to identify key influencing factors for Treatment outcomes while quantifying their weights, establish Critical Thresholds for core indicators, provide scientific evidence for individualized Therapeutic regimen development in CHH, and offer methodological references for precision medicine research in Rare Endocrine Diseases.
2 Materials and methods2.1 Study participantsThis single-center retrospective cohort study enrolled adolescent male patients with CHH who were diagnosed and received fertility induction therapy at our hospital’s endocrinology department between January 2010 and October 2025.
2.1.1 Inclusion CriteriaReferencing our research group’s previously published article (Wang et al., 2017), the specific criteria were:
Advanced age group: 1. Male >14 years old without pubertal development, defined as testicular volume <4 mL; 2. Bone age ≥12 years 3. Baseline hormone levels indicating prepubertal status (serum androgen levels ≤20 ng/dL in males); 4. No elevation in baseline gonadotropin levels; 5. Absence of thyroid or growth hormone axis abnormalities; 6. Normal karyotype; 7. Developmental anomalies such as olfactory bulbs or olfactory tracts may be observed in the hypothalamic and pituitary regions, but no space-occupying lesions are present; 8. May be accompanied by micropenis, cryptorchidism, or hypospadias;
Younger age group: 1. Males >10 years old but ≤14 years old; 2. Meets criteria 3 to 8 of the advanced age group; 3. Positive family history of CHH; 4. According to ACMG standards, genetic testing reveals pathogenic or likely pathogenic mutations in candidate genes for CHH.
Patients without pubertal progression: 1. Testicular volume >4 mL, or basal testosterone level >20 ng/dL, or LH/FSH levels reaching early pubertal levels; 2. Presence of olfactory abnormalities or imaging-confirmed dysplasia of olfactory bulbs/tracts; 3. Adolescents demonstrating no pubertal progression during ≥6 months of clinical follow-up (Puberty Arrest) meeting definitive diagnostic criteria for Kallmann Syndrome were also enrolled.
2.1.2 Exclusion criteria① Identifiable etiologies (e.g., confirmed chromosomal abnormalities, trauma, surgery) or other known disorders causing hypogonadism (e.g., Prader-Willi Syndrome with sexual infantilism); ② Chronic systemic diseases (e.g., uremia, thalassemia, poorly controlled diabetes); ③ Protein-energy malnutrition; ④ Eating disorders (e.g., anorexia nervosa, hyperphagia); ⑤ Intracranial space-occupying lesions (e.g., pituitary tumors or post-operative status); ⑥ Absence of fertility-inducing treatment (GnRH pump or gonadotropin therapy); ⑦ Lost to follow-up within 6 months of treatment initiation.
2.1.3 Treatment outcomesNocturnal emission group: Fertility induction therapy for at least 6 months, during which baseline Testosterone progressively reached >200 ng/dL, with testis and penis enlargement, culminating in nocturnal emission confirmed by patient or guardian report and documented by the follow-up physician during regular outpatient interviews (typically occurring within 12–24 months, though extended treatment duration was required in some cases).
Non-ejaculation group: Fertility induction therapy for at least 6 months, during which baseline testosterone remained <200 ng/dL with minimal testis/penis growth; or cases exhibiting initial testosterone elevation (>200 ng/dL) with genital development, but subsequently declined to <200 ng/dL with arrested/reversed genital growth, leading to therapy discontinuation and transition to testosterone replacement therapy.
Ultimately, 65 patients were included and categorized into the nocturnal emission group (55 cases) and non-ejaculation group (10 cases) based on treatment outcomes.
2.2 Data sources and indicator classification2.2.1 Data collectionResearch data were sourced from the hospital’s electronic medical record system, laboratory information management system, and genetic testing database. Baseline patient information was independently extracted by two endocrinologists and cross-verified to ensure data accuracy. Penile length was measured in the flaccid state as the stretched length from the pubic bone to the tip of the glans, using a rigid ruler pressed firmly against the pubic symphysis to compress the suprapubic fat pad. Penile diameter was measured at the mid-shaft. All measurements were performed by experienced pediatric endocrinologists.
2.2.2 Indicator systemA total of 15 pre-treatment baseline indicators across 4 major categories were collected as follows: ① Medical history: treatment initiation age, history of cryptorchidism, family history of constitutional delayed development or menstrual irregularities, disease type, presence of dual CHH (defined as testosterone levels <100 ng/dL following standard or extended human chorionic gonadotropin (HCG) stimulation testing), and treatment modality. ② Physical examination: testicular volume (measured using Prader orchidometer), penile diameter, penile length; ③ Serological indicators: luteinizing hormone (LH), FSH, testosterone, AMH, INH-b; ③ Genetics: genetic testing for phenotype-associated gene variants.
2.3 Therapeutic methodsFertility induction therapy consisted of either gonadotropin-releasing hormone (GnRH) pulsatile pump therapy or combined gonadotropin therapy. The GnRH pulsatile pump regimen involved subcutaneous infusion of 5–12 µg per pulse every 90 min, delivering 16 pulses per 24 h. Initial dosing was conservative, with subsequent titration based on Serum Testosterone Levels, targeting a concentration of 200–500 ng/dL. Combined gonadotropin therapy employed two regimens: Protocol 1 involved initial monotherapy with hCG (1000–2000 IU) administered intramuscularly every other day or twice weekly. Upon achieving a testosterone level of 200 ng/dL, human menopausal gonadotropin (hMG) 75 IU was added weekly; protocol 2 utilized combined administration of hCG (500–2000 IU) and hMG (75–150 IU) administered 1–3 times weekly. Throughout treatment, hCG doses were adjusted to maintain Testosterone within target range.
All patients underwent follow-up at months 1 and 3 post-treatment initiation, subsequently every 3 months, encompassing hormone level assessments and physical examinations. Serum testosterone Levels were monitored 12–24 h post-hCG/HMG administration or during continuous GnRH pump therapy. Testicular volume was assessed using Prader orchidometer and ultrasonography, with bilateral mean values used for analysis.
2.4 Research instrumentsAll statistical analyses and machine learning modeling were conducted on MATLAB R2024b, implementing data cleansing, stratified cross-validation, Synthetic Minority Over-sampling Technique (SMOTE), Grid Search, and random forest training to ensure reproducible results.
2.5 Research methodsFigure 1 illustrates the research methodology workflow. This study employed a retrospective case-control design to develop and validate a machine learning algorithm-based predictive model for treatment outcomes in CHH. The research methodology was divided into six phases: Initially, Inclusion Criteria for study subjects were established, and subjects were retrieved from the electronic health record (EHR) system. Clinical data then underwent standardized preprocessing, including missing value imputation (median imputation for continuous variables, mode imputation for categorical variables) and low-variance feature filtering. Subsequently, multi-dimensional feature screening was conducted through univariate statistical analysis (Mann-Whitney U test for continuous variables, Chi-square test for categorical variables) combined with random forest feature importance ranking to identify key predictors. To address the Class Imbalance Problem, SMOTE was utilized to enhance the model’s recognition capability for the minority class. During the core modeling phase, the random forest algorithm was employed, and model performance was evaluated through 5-fold stratified cross-validation. Optimal prediction thresholds were determined through ROC curve analysis and Youden index calculation, with specific threshold recommendations provided for different clinical scenarios (avoiding missed diagnosis, avoiding overtreatment, and routine diagnosis). Through systematic comparison with control models including SVM, Logistic Regression, and k-NN, the predictive efficacy and clinical applicability of the developed model were comprehensively evaluated. Ultimately, the influencing factors of treatment outcomes were investigated through feature importance analysis based on machine learning.

Flowchart of research methodology.
2.6 Data preprocessingMissing value handling: Continuous indicators (e.g., LH, FSH levels) were imputed using the median of their respective groups (success/failure groups); Categorical indicators (e.g., genetic mutation status) were imputed using the mode. Ultimately, the proportion of missing data in all cases was <5%.
Outlier identification and handling: Outliers were identified using the boxplot method (IQR±1.5). Following clinical expertise assessment, non-clinically significant outliers were truncated, while clinically specific outliers were retained and annotated.
Low-variance Feature Filtering: Core function is to eliminate features with minimal numerical fluctuation and no outcome discriminability (e.g., homogeneous test indicators). By setting a 0.01 variance threshold, high-information features were screened while guaranteeing retention of minimum indicators. This approach reduced computational load, mitigated overfitting, and allowed subsequent modeling to focus on clinically discriminative indicators, thereby establishing the foundation for the precision and generalizability of the CHH treatment outcome prediction model.
Class Imbalance Correction: Due to differences in sample sizes between groups (success group: 55 cases, failure group: 14 cases), SMOTE was employed for oversampling to balance the training set sample distribution; Simultaneously, class weights were set during model training to mitigate the impact of sample imbalance on the model.
Data Standardization and Encoding: Continuous indicators were standardized using Z-score normalization; binary indicators (e.g., presence of genetic mutations) were encoded using 0–1 binary encoding.
2.7 Feature selectionThe core objective of the feature selection step is to identify core features strongly associated with treatment outcomes (success/failure) and possessing high predictive value from preprocessed high-variance features. This serves to reduce model computational complexity, prevent overfitting, and enhance the clinical interpretability of the model. This step employs a dual logic of ‘preliminary screening via statistical tests + rigorous screening based on model importance,’ dynamically adapted to the characteristics of the clinical data. The specific workflow and its significance are outlined below:
First, univariate statistical tests were conducted for preliminary screening. Considering the differences between continuous (e.g., hormone levels) and categorical (e.g., therapeutic regimen) features in clinical data, appropriate methods were employed: Mann-Whitney U test (non-parametric test, given the non-normal distribution characteristics of clinical data) for continuous features and Chi-square test (to analyze feature-outcome associations) for categorical features. The P-values indicating associations between each feature and the outcome were calculated, with features demonstrating P < 0.05 selected as significant during the initial screening. Features lacking discriminative power (e.g., single-value features) were directly marked as having no value (P = 1) to ensure the statistical validity of the preliminary screening results.
Secondly, adaptive dimensional optimization was implemented by setting the maximum feature count. Considering the typically limited sample size in clinical research, the maximum feature count was dynamically adjusted based on sample size constraints: ≤50 samples used 1/5 features, ≤100 samples used 1/4 features, otherwise 1/3 features were utilized. Concurrently, upper and lower thresholds (2–20 features) were constrained to prevent overfitting caused by ‘high-dimensional data with limited samples,’ thereby balancing feature dimensions with model generalization capability to accommodate the scarcity characteristic of CHH clinical data.
Finally, feature importance-based rigorous screening was performed using the random forest algorithm. Temporary random forest models were trained to quantify feature importance through ‘Out-of-Bag Error Change,’ with the top 80% of high-importance features being selected; If fewer than 2 features remain post-screening, the fallback mechanism retains the optimal feature from initial screening to ensure the feature set meets minimum validity thresholds. The final output specifies both the number of retained features and their identities, explicitly identifying core predictive targets for CHH treatment outcomes. Through the above comprehensive workflow, 7 core features were ultimately selected from the original 15 baseline indicators for training the final random forest prediction model. This process corresponds to the ‘Feature Screening’ stage in Figure 1.
This comprehensive workflow balances statistical significance with predictive value by excluding irrelevant features and focusing on core clinical indicators (e.g., key hormone levels, treatment phase parameters). This approach enhances subsequent model training efficiency and generalizability while providing clinically actionable targets for decision-making, thereby strengthening the model’s clinical utility.
2.8 Development and training of machine learning models2.8.1 Model selectionRandom forest model was selected as the primary methodology for this investigation. The random forest algorithm, as an ensemble tree-based model, accommodates nonlinear relationships, adapts to small-sample datasets, and outputs Feature Importance metrics, thereby aligning with the core requirements of this study. Through dual perturbation mechanisms—bootstrapping aggregation and feature random subspace—the random forest architecture inherently suppresses overfitting in our limited cohort of 65 CHH patients; Its robust handling of mixed data types, missing values, and nonlinear associations obviates stringent distributional assumptions and complex preprocessing. The ensemble voting mechanism seamlessly generates class probabilities, facilitating subsequent optimization of clinical decision thresholds; The built-in Out-of-Bag (OOB) error and Feature Importance metrics provide directly interpretable evidence for dimensionality reduction under lenient significance thresholds, balancing predictive performance with mechanistic insights.
2.8.2 Stratified cross-validation designIn CHH clinical treatment outcome prediction studies, the core design of stratified k-fold cross-validation (k = 5) addresses both clinical data characteristics and the rigor requirements of clinical research.
Resolving the class imbalance problem in clinical data to ensure unbiased evaluation. In CHH clinical practice, ‘treatment failure’ samples are typically significantly fewer than ‘success’ samples. Conventional cross-validation may suffer from imbalanced sample distribution in individual folds (e.g., absence of failure samples in a particular fold), causing models to exhibit bias toward predicting the majority class. The stratified design ensures consistent class distribution across all folds by performing stratified sampling based on the ‘success/failure’ ratio of the original data. This approach enables the model to learn features of both outcome categories in each fold, preventing distorted evaluation due to sampling bias. Crucially, it guarantees objective assessment of predictive capability for clinically critical ‘failure outcomes’.
Maximizes the utilization of limited clinical samples to enhance evaluation reliability. Given the difficulty in obtaining clinical research samples and their limited quantity, stratified cross-validation employs an iterative ‘training-testing’ cycle design. This allows every sample to participate in both training and testing phases, maximizing information extraction from the dataset. Compared to single train-test splits, this approach significantly reduces evaluation error caused by random partitioning. The selection of 5-fold cross-validation balances evaluation efficiency and result stability, mitigating the risk of high variance from insufficient folds while avoiding computational redundancy from excessive folds. This approach ensures that core metrics (mean ± SD), including AUC and F1-score, more accurately reflect the model’s true generalization capability.
This design aligns with the stringent reliability requirements of clinical decision-making, thereby enhancing model trustworthiness. Clinical outcome prediction directly informs personalized treatment regimen development, making model stability of paramount importance. Stratified cross-validation demonstrates model performance consistency across distinct sample subsets by revealing metric fluctuations (e.g., AUC standard deviation) throughout multiple validation folds. Minimal fluctuations in metrics across folds indicate that the model maintains stable predictive capabilities for CHH patients with varying feature distributions, providing crucial reliability evidence for clinical implementation; this prevents clinical decision-making risks arising from model failure in specific patient subgroups. Furthermore, it reinforces the rigor of research conclusions and meets requirements for both academic validation and clinical translation. Stratified cross-validation represents the standardized evaluation methodology in clinical machine learning research, offering strong reproducibility that effectively verifies the model’s generalization capability is not attributable to chance sample partitioning. The multidimensional indicators (Accuracy, Precision, Recall, AUC) and their fluctuation ranges obtained through this study design provide robust quantitative support for the conclusion that ‘the model can effectively predict CHH treatment outcomes,’ enhancing the academic persuasiveness of the research findings and establishing a methodological foundation for subsequent clinical translation of the model.
2.8.3 SMOTE oversamplingIn this study, addressing the extremely imbalanced scenario with only 65 CHH cases and 10 positive events, SMOTE was employed to infuse ‘synthetic yet credible’ minority class samples into the random forest through feature-space interpolation, fundamentally mitigating model bias induced by class imbalance. The algorithm generates new points along the difference vectors between randomly selected minority class samples and their k-NN, preserving the original feature distribution while avoiding overfitting caused by simple duplication. Continuous variables and discrete variables were processed using linear interpolation and mode voting respectively, ensuring logical consistency across blood test results, physical examination indicators, and genetic history. After SMOTE balancing, the success-to-failure ratio in the training set was adjusted from 4.9:1 to approximately 1:1. This enabled random forest’s bootstrap sampling to be adequately exposed to failure samples during node splitting in each decision tree, thereby correctly learning the decision boundary and significantly improving Recall and F1-score. Meanwhile, the introduction of synthetic samples increases the coverage density in the feature space and reduces the sensitivity of decision tree partitioning to noise. Combined with random forest’s out-of-bag error estimation, this approach could further suppress the risk of overfitting. Notably, SMOTE operates exclusively on the training set, while the validation and test sets retain their original distributions. Consequently, performance metrics can still authentically reflect the model’s generalization capability in real clinical scenarios. For rare diseases such as CHH with extremely high data acquisition costs, SMOTE employs a low-cost ‘data augmentation’ strategy. This enables random forest to achieve statistical power commensurate with the sample size without increasing patient burden, thereby providing a robust foundation for subsequent feature screening, threshold optimization, and clinical interpretation.
2.8.4 Hyperparameter tuningThis code implementation does not include explicit hyperparameter optimization processes (such as grid search, Bayesian Optimization, or other iterative optimization methods). Instead, it employs a core strategy of ‘self-adaptive parameter adjustment + empirical value setting’ to accommodate the characteristics of CHH clinical data. Firstly, key hyperparameters (e.g., the number of trees in random forest, cross-validation folds, SMOTE neighbor count) were predetermined with empirical values through a structured parameter set, without comparative screening of multiple parameter combinations. Secondly, model parameters are dynamically adjusted according to data scale: in random forest training, the ‘number of features per tree’ (max_features) is adaptively determined as the ‘square root of feature count’, while the ‘minimum leaf node size’ (min_samples_leaf) is adjusted based on total sample size (not less than 2 and set to 1/20th of the sample count). During the feature selection phase, the maximum feature count was dynamically determined based on sample size (≤50 samples: 1/5 features, ≤100 samples: 1/4 features, otherwise: 1/3 features) to balance dimensions and generalizability. Although not explicitly optimized, this design accommodates the limited availability of clinical samples. By integrating empirical values with data adaptability, it maintains model stability while simplifying computational workflows, effectively mitigating overfitting risks associated with hyperparameter tuning in small sample sizes.
2.9 Feature importance evaluation methodologyThe random forest model developed in this study, constructed through a comprehensive workflow of ‘data preprocessing-feature screening-class balancing-stratified validation’, demonstrated excellent and stable predictive performance, indicating substantial potential for clinical application. The core strength of the model lies in its optimal adaptation to clinical data characteristics. By adopting an implicit hyperparameter tuning strategy, it streamlines the workflow while ensuring reliability, with specific manifestations and values as follows:
Feature importance assessment within our random forest model serves as a pivotal step for screening core predictive indicators of CHH treatment outcomes and enhancing model interpretability. Centered on ‘OOB Permuted Predictor Delta Error’ to quantify feature value, this approach permeates the entire workflow from feature selection to result visualization, aligning with both clinical data characteristics and modeling requirements.
Regarding evaluation principles and implementation logic, the model validates Feature Importance through ‘OOB’ data: For each feature, its values in OOB data are randomly permuted, followed by recalculation of the model’s out-of-bag prediction error; The feature importance value is calculated as the difference between post-permutation error and original error. A larger difference indicates a greater contribution of the feature to model prediction, and its absence would cause significant degradation in model performance. This method requires no additional validation set, adapting to the limited availability of clinical samples, while effectively circumventing interference from overfitting in evaluation results.
The application in the modeling workflow comprises two core phases: First, rigorous screening during feature selection. Based on preprocessed high-variance features, we train a provisional random forest comprising 50 decision trees to compute feature importance. Features are ranked in descending order of importance, with the top 80% high-importance features selected. A safeguard mechanism ensures retention of at least two features. This approach compensates for the limitation of univariate statistical tests that focus solely on pairwise ‘feature-outcome’ associations, further eliminating redundant features from the perspective of model prediction logic. Secondly, the visualization of the final model involved training the definitive random forest on the complete balanced dataset (post-SMOTE). Feature Importance values were extracted and presented in a bar chart, clearly annotating the predictive weights of clinical characteristics (e.g., hormone levels, treatment phase indicators) to provide an intuitive basis for clinical interpretation.
The core value of this evaluation method lies in striking a balance between modeling performance and clinical significance: on one hand, by screening high-importance features, it reduces model computational complexity and enhances predictive stability, avoiding overfitting caused by irrelevant features; On the other hand, the quantified Feature Importance results delineate core therapeutic targets for predicting CHH treatment outcomes, transforming the model from a ‘black box’ into an interpretable clinical tool. This empowers clinicians to identify key indicators influencing therapeutic efficacy, provides data-driven support for formulating personalized treatment regimens, and enhances the clinical translational value of research findings.
2.10 Model performance evaluation metricsThis study employs multi-dimensional metrics to comprehensively evaluate the random forest model’s performance in predicting clinical treatment outcomes for CHH. Core metrics include Accuracy, Precision, Recall, F1 score, and AUC value, specifically adapted to address class imbalance characteristics of clinical data to ensure objective and practical evaluation. Accuracy reflects the model’s overall prediction correctness, serving as a fundamental evaluation benchmark; precision measures the proportion of true positives among samples predicted as ‘failure to achieve nocturnal emission post-treatment,’ preventing excessive misjudgment of unsuccessful cases. Recall focuses on the detection capability for genuinely unsuccessful cases, aligning with the clinical core need for early warning of adverse outcomes. The F1 score is the harmonic mean of precision and recall, balancing the trade-off between these two metrics and adapting well to the evaluation of imbalanced datasets. The AUC value, calculated via the ROC curve, comprehensively reflects the model’s ability to distinguish between ‘treatment success’ and ‘treatment failure,’ offering enhanced stability. All metrics were calculated based on a 5-fold stratified cross-validation, reporting the mean ± standard deviation. This approach mitigates the randomness of single evaluations, quantifies model variability, and provides robust quantitative support for the reliability of the model’s clinical application.
3 ResultsAs illustrated in Figure 2, the model demonstrates exceptional comprehensive performance and robust stability in training outcomes. Core evaluation metrics based on 5-fold stratified cross-validation exhibited outstanding performance: Accuracy reached 0.84 ± 0.12, indicating a high level of overall predictive correctness; precision was 0.82 ± 0.13, signifying a high proportion of true positives among cases predicted as ‘failure to achieve nocturnal emission post-treatment’, which can effectively mitigate clinical decision-making risks associated with excessive misjudgment of treatment failures; recall of 0.89 ± 0.08 reflects the model’s strong detection capability for genuine failure samples, aligning with the clinical imperative for early warning of adverse outcomes; F1 score 0.85 ± 0.10 indicates balanced coordination between Precision and Recall, demonstrating adaptability to the class imbalance characteristic in clinical data; the particularly critical AUC value reached 0.95 ± 0.04, signifying the model’s exceptional comprehensive ability to differentiate between ‘successful nocturnal emission post-treatment’ and ‘failure to achieve nocturnal emission post-treatment’ outcomes, with minimal standard deviation reflecting high consistency in model performance across different sample subsets.

Random Forest model 5-fold stratified cross-validation.
The stability of model performance benefits from scientific workflow design: balancing Minority class (failure samples) distribution via SMOTE oversampling prevents model bias toward Majority class; stratified cross-validation ensures consistent sample distribution across all folds with the overall dataset, enhancing evaluation reliability; during the feature selection phase, core clinical indicators were screened by combining statistical testing with random forest importance assessment, simultaneously simplifying model architecture and enhancing interpretability. In summary, this model not only demonstrates predictive advantages of high accuracy and strong discriminative capability but also reduces the barrier to application by adapting to clinical data characteristics. It empowers clinicians to precisely identify key factors influencing treatment outcomes and provides reliable data support for formulating personalized treatment regimens for CHH patients.
3.1 Quantitative ranking of influencing factors for treatment outcomes3.1.1 Inherent feature importance results of the modelAs shown in Figure 3, the feature importance ranking of the random forest model is as follows: Cryptorchidism (Out-of-Bag Error Change 1.2597), Pre-Treatment Penile Length (0.7341), Pre-Treatment Penis Diameter (0.7295), Pre-Treatment FSH (0.6145), Pre-Treatment AMH (0.5902), Pre-Treatment Mean Testicular Volume (0.5571), Pre-Treatment LH (0.4258). Notably, the feature importance value of ‘history of cryptorchidism’ (1.2597) as a binary variable is significantly higher than that of other continuous variables, highlighting its core position in the predictive model.

Model feature importance ranking.
3.1.2 Determination of critical indicator thresholdsThis study developed a CHH treatment outcomes prediction model through ROC curve analysis, identifying the optimal critical threshold and core predictive features to provide a quantitative basis for formulating personalized clinical treatment strategies. The results are analyzed as follows.
Critical threshold analysis revealed (Figure 4) that based on the principle of maximizing the Youden index, the optimal threshold across model folds was 0.601 ± 0.137. This threshold achieved a treatment failure identification rate (TPR) of 0.89 with only a 0.05 misclassification rate (FPR) for successfully treated nocturnal emission cases, yielding a Youden index of 0.836. These results demonstrate the model’s outstanding predictive accuracy and clinical utility. To address distinct clinical management needs, as Table 1 shows, this study proposes a stratified threshold recommendation strategy: Scenario 1 (prioritizing avoidance of false negatives in treatment-nonresponsive patients) adopts a 0.501 threshold, ensuring an 87% failure sample identification rate at the cost of a marginally elevated false positive rate (0.10), applicable for high-risk population screening; Scenario 2 (prioritizing avoidance of overtreatment) employs a 0.701 threshold, balancing medical resource wastage while maintaining an 84% failure sample identification rate, suitable for conservative management of low-risk patients; Scenario 3 (standard clinical management) utilizes the optimal 0.601 threshold, achieving precise equilibrium between risk of false negatives and overtreatment.

Roc curve of the Model’s optimal threshold.
ScenarioThresholdTrue positive rate (TPR)False positive rate (FPR)Prioritizing avoidance of false negatives in treatment-nonresponsive patients0.5010.870.16Prioritizing avoidance of overtreatment0.7010.860.22Standard clinical management0.601 (Youden optimal threshold)0.890.05Recommended clinical thresholds for CHH treatment prediction.
The stratified cutoff thresholds established in this study can be tailored to different clinical scenario needs. Cryptorchidism, pretreatment external genitalia development, and hormone-related indicators (including AMH, INH-b, and FSH) serve as core predictive factors for CHH treatment outcomes. Their synergistic application provides critical references for treatment risk assessment and personalized regimen optimization in CHH patients.
3.2 Comparison with other machine learning methodsThis study conducted comparisons with three classical machine learning models—SVM, Logistic Regression, and k-NN—by replacing the random forest model in the research procedure. As illustrated in Figure 5a, the model comparison results demonstrate that the random forest model (AUC = 0.9504) achieved optimal performance in predicting CHH treatment outcomes, with its AUC value being significantly higher than those of other models (an 11.29% improvement), showcasing robust discriminatory capacity. The ROC curve comparison graph visually demonstrates the classification performance of each model: The random forest curve is closest to the top-left corner, followed by k-NN, while the SVM and Logistic Regression curves are nearer to the diagonal, indicating limited discriminative capability. Notably, while k-NN slightly underperforms random forest in AUC (0.8926 vs. 0.9504), its Recall reaches 1.0, achieving complete identification of cases without nocturnal emission. This characteristic positions it as an effective supplementary tool for high-risk screening.

Performance comparison of different machine learning methods: (a) Comparative ROC Curves of Models. (b) Bar chart of Model Performance Differences. (c) Radar Chart of Model Performance Disparities.
The bar chart in Figure 5b reveals differences in model performance across multiple dimensions. Random forest demonstrated superior comprehensive performance and computational efficiency, leading in Accuracy (0.8182), F1 score (0.8333), and AUC metrics. While achieving exceptional Recall (1.0), k-NN’s performance came at the cost of slightly lower Precision, yet provides significant value for clinical early-warning systems. SVM and Logistic Regression showed k-NN balanced but overall lower performance across all metrics, indicating the limitations of Linear Models when handling this non-linear clinical data.
Figure 5c radar chart geometrically visualizes model characteristics: random forest displays a near-pentagonal balanced profile, demonstrating robustness across all evaluation dimensions; The curve demonstrated significant convexity along the recall axis, exhibiting an asymmetric profile indicative of high sensitivity yet relatively insufficient specificity; the compact contours of SVM and Logistic Regression confirmed their limited performance. Comprehensive visualization analyses revealed that random forest exhibited optimal balance, discriminative power, and practicality for CHH treatment prediction, establishing it as the preferred model for clinical decision support. Meanwhile, k-NN’s high-recall characteristic renders it suitable for auxiliary screening of high-risk cases.
4 DiscussionThis study successfully achieved high-performance prediction of fertility induction therapy outcomes (marked by nocturnal emission) in adolescent male patients with CHH by constructing a predictive model based on the random forest algorithm. The model demonstrated exceptional d
Comments (0)