Within-sibling attenuation of polygenic risk score accuracy: investigating the effects of principal component analysis, LD score regression, and mixed model association in the UK Biobank

Datasets and traits analyzed

Analyses were conducted using the 2018 UK Biobank genotype release in the self-reported White ancestry subset. Principal component analysis (PCA) plots are shown in Supplementary Figures S1 and S2. We examined four disease traits for which polygenic risk scores have shown potential clinical utility (Sudlow et al. 2015; Bycroft et al. 2018). For the purposes of comparison, we examined the effects of controlling for population stratification on a binarized version of the educational attainment phenotype which has been well-established as having issues with confounding and within-family attenuation (Lee et al. 2018; Selzam et al. 2019; Lello et al. 2020). The relative importance of contributing factors to this attenuation is not currently settled, with some suggestion that population stratification may play a major role (Tan et al. 2024; Smith et al. 2025). Other authors have also found a large role for indirect genetic effects in this trait (Mostafavi et al. 2020; Howe et al. 2022).

Summary statistics from each GWAS were used to build PRS models used to classify between discordant pairs (consisting of one case and one control in each pair). Two test sets were evaluated, one consisting of sibling pairs and the other of population-level unrelated pairs (non-siblings). To account for the effects of different levels of polygenicity, three SNP-set sizes were examined during PRS construction (1 K, 10 K, and 100 K). The total sample size of each phenotype, and the discordant sibling and non-sibling test sets can be seen in Tables 1 and 2 respectively.

Table 1 Sample sizes in training GWAS for positive and negative phenotype classes in the UK BiobankTable 2 The number of individuals in discordant sibling test sets from the UK Biobank

The primary outcome was the attenuation in predictive performance observed when moving from population-level to within-sibling prediction.

Baseline PRS accuracy and attenuation

To establish baseline performance metrics and quantify the degree of confounding present prior to population structure adjustment, we first constructed PRSs using standard GLM-based GWAS without controlling for population stratification (i.e. no PCs were included as covariates). The predictive accuracy of these PRSs in both the population-level and within-sibling settings is presented in Table 3. Overall, predictive accuracy was modest across all traits, with no PRS exceeding 65% classification accuracy in the population-level setting. As expected, we observed reduced accuracy among discordant sibling pairs compared to unrelated individuals across most trait-SNP set combinations, reflecting within-sibling attenuation. Educational attainment exhibited the largest attenuation, with the majority of classification accuracy lost when transitioning from the population-level to within-sibling setting, consistent with prior literature documenting substantial confounding in this trait (Lee et al. 2018; Selzam et al. 2019; Lello et al. 2020).

Table 3 Baseline classification accuracy and attenuation for standard GWAS-based PRS performance on sibling and population-level (non-sibling) discordant pairs in the UK Biobank

Among the disease traits, attenuation was generally significant but more moderate in magnitude. A notable exception to this was that the 1000 SNP prostate cancer model showed negative attenuation, with the PRS demonstrating higher predictive accuracy in the sibling set than among unrelated individuals. This pattern was not observed for other SNP set sizes and may be reflective of stochastic variation.

We did not observe a consistent relationship between SNP set size and either predictive accuracy or attenuation magnitude. Specifically, larger SNP sets did not consistently yield improved prediction in the population-level setting, nor did they exhibit increased attenuation as might be expected if the inclusion of SNPs with lower P-values was more susceptible to capturing confounding effects. Based on performance in the population-level test set, we selected the best-performing SNP set size for each trait to carry forward in subsequent analyses examining the impact of PC adjustment (bolded in Table 3).

Changes in the LDSC intercept with increasing PC inclusion

To evaluate whether standard population structure adjustment methods effectively reduce confounding as traditionally measured, we examined the effect of including an increasing number of PCs on the LDSC intercept and ratio across all traits. The LDSC intercept results are presented in Fig. 1, with corresponding LDSC ratio trends shown in Supplementary Figure S4. The intercept decreased for most traits as the number of PCs included in the training GWAS increased, though the magnitude and pattern of reduction varied substantially across phenotypes.

Fig. 1Fig. 1

Changes in the LD score regression (LDSC) intercept with increasing inclusion of principal components (PCs) across phenotypes: coronary artery disease (CAD), type 2 diabetes (T2D), breast cancer (BrC), prostate cancer (PrC), and binarized educational attainment (EA). LDSC summary statistics were derived from GLM-based GWAS results. Shaded regions represent standard errors. The y-axis limit is higher for educational attainment due to greater overall inflation

Consistent with the baseline attenuation results, educational attainment exhibited the strongest response to PC-adjustment. Both the intercept and ratio were substantially elevated at baseline (no PCs included), indicating considerable confounding. While subsequent inclusion of PCs produced marked reductions in both metrics, residual confounding likely remained even when all sixteen principal components were included as covariates (Wang et al. 2023). Coronary artery disease and type 2 diabetes demonstrated moderate baseline evidence of confounding, with intercepts modestly elevated above 1.0. Both traits showed downward trends in the intercept as more PCs were included, though the absolute magnitude of change was not extreme. Notably, after PC correction with the full complement of 16 PCs from data field 22009, both traits achieved LDSC ratios below 1.1, falling under a rule-of-thumb threshold for minor confounding (Bulik-Sullivan Lab 2025). Breast and prostate cancer exhibited low absolute levels of confounding at baseline. The LDSC ratios for these traits were more moderate but associated with large standard errors, possibly reflecting sample size limitations for these two phenotypes. Finally, similar patterns as these trends were observed for the genomic inflation factor λ across all traits (Supplementary Figure S3).

PRS accuracy across levels of PC inclusion

To assess whether reductions in test statistic inflation translate to changes in predictive performance, we constructed PRSs for each phenotype across all levels of PC inclusion, with PCs included as covariates in both the training GWAS and during PRS prediction.

Results for the best-performing SNP set size for each trait are presented in Fig. 2. Among unrelated individuals, we observed no consistent monotonic relationship between the number of included PCs and predictive accuracy, though modest variation across PC levels was evident. Among discordant sibling pairs, there was a slight overall trend toward improved performance with increasing PC inclusion, though this pattern was not universal across all traits.

Fig. 2Fig. 2

PRS classification accuracy across levels of PC inclusion for sibling and non-sibling discordant pairs, shown for the best-performing SNP-set size per phenotype

With respect to phenotype-specific patterns, the largest variation in accuracy was observed for breast and prostate cancer. Across both sibling and population pairs, the first six PCs generally increased the accuracy for the EA phenotype, while only modest changes in magnitude were observed for CAD and T2D. When examining the full range of SNP set sizes, the absence of a consistent trend in either unrelated individuals or sibling pairs became even more apparent (Supplementary Figures S5–S9).

The principal components used in these analyses were computed across the entire UK Biobank cohort, including individuals of more diverse ancestry than those included in our ancestry-filtered GWAS. To assess whether this might diminish the effectiveness of PC-adjustment for our analysis cohort, we recomputed PCs using only individuals from the post-QC, self-described White subset. When these cohort-specific PCs were used in GWAS and PRS construction, they did not substantially alter predictive accuracy for any of the disease traits examined (Supplementary Tables S2 and S3).

Whilst overall predictive accuracy is informative, the primary outcome of interest for evaluating confounding is not absolute performance. PRS accuracy may reflect both causal genetic effects and confounding, whereas within-sibling attenuation directly quantifies the degree to which population-level associations fail to replicate in a setting that controls for both genetic background and environmental confounding.

Predictive attenuation across levels of PC inclusion

We next examined whether PC-adjustment reduces within-sibling attenuation. Reductions in attenuation when transitioning from the population-level to within-sibling setting would indicate that PC-adjustment successfully controls for environmental and genetic background confounding. The experimental design here allows us to isolate the specific effects of environmental and genetic background confounding that PC-adjustment is intended to target. Whilst some persistent attenuation is expected due to shared genetics between siblings and the other sources of confounding not addressed by PC-adjustment, reductions in attenuation with increasing PC inclusion can provide evidence that standard population structure adjustment methods are successfully mitigating the confounding they are specifically designed to control.

Attenuation results for each trait are presented in Fig. 3. Across all traits, we did not observe a consistent decrease in attenuation as more PCs were included. A weak trend toward reduced attenuation may be present in some cases, although the relationship was neither monotonic nor uniform. Notably, attenuation sometimes increased or decreased in a non-systematic manner depending on the specific number of included PCs, suggesting that the optimal number of PCs for minimizing attenuation is not predictable in advance of within-sibling validation, or through observation of the LDSC intercept. This lack of consistent pattern was further confirmed when examining all SNP set sizes (Supplementary Figure S10).

Fig. 3Fig. 3

Attenuation in PRS predictive performance between sibling and non-sibling discordant pairs across levels of PC inclusion, shown for the best-performing SNP-set size per phenotype

Educational attainment exhibited persistently high attenuation across all levels of PC inclusion, despite the substantial reduction in LDSC intercept documented above. Similarly, for coronary artery disease and type 2 diabetes, the decreases in LDSC intercept observed with increasing PC inclusion were not accompanied by corresponding reductions in within-sibling attenuation. These findings indicate that strong-to-moderate reductions in confounding as measured by the LDSC intercept do not translate to meaningful improvements in within-sibling predictive validity. This suggests some limitations in using test statistic inflation metrics as proxies for the causal validity of polygenic risk scores.

Our results are consistent with prior work suggesting that population structure (as captured by PCs) may contribute only modestly to PRS performance for some traits in the UK Biobank. For example, Lello et al. showed for height that regressing the phenotype on principal components yields low explanatory power, implying limited variance attributable to PC-defined structure (Lello et al. 2018). Similarly, although several PCs were significantly associated with each phenotype in our data, we did not observe a clear relationship between the significant PCs and specific changes in predictive performance or within-sibling attenuation, despite the downward trends in overall inflation mediated by PC-adjustment (Supplementary Table S1).

Additional effects of mixed model GWAS

To investigate whether GRM-based mixed models provide additional control over confounding beyond PC-adjustment alone, we performed GWAS using fastGWA-GLMM with the full set of 16 PCs included as covariates. We compared the resulting PRSs to those derived from standard GLM-based GWAS with either 0 PCs or 16 PCs.

Results for the best-performing SNP set size for each trait are presented in Table 4. As already shown in the previous findings, inclusion of the full set of PCs may have produced slight improvements in accuracy and reductions in attenuation relative to models with no PC adjustment. However, the addition of mixed model correction provided no apparent benefit beyond that achieved through PC inclusion alone. Within-sibling attenuation was nearly identical between the 16-PC GLM and GLMM + 16-PC approaches across most traits examined.

Table 4 Classification accuracy and within-sibling attenuation for PRS derived from GWAS across population structure adjustment methods

This pattern was also largely consistent across all SNP set sizes (Tables 3, S5, and S6). Furthermore, the GLMM approach did not meaningfully reduce the genomic inflation factor λ beyond the level achieved with 16 PCs in a standard GLM framework (Supplementary Table S4). Taken together, these results suggest that in large and relatively homogeneous cohorts such as the UK Biobank, mixed model methods do not offer substantive improvements over PC-adjustment for controlling confounding effects attributable to population stratification.

To more formally assess the factors influencing classification accuracy, we fit a generalized linear mixed model (GLMM) with phenotype, SNP-set size, pair type (sibling vs. non-sibling), and adjustment method (0-PC, 16-PC, and GLMM + 16-PC) as fixed effects. This analysis allowed the relative importance of each factor to be evaluated while accounting for repeated measurements across traits and models.

Phenotype emerged as the strongest predictor of classification accuracy (χ2 = 114, p < 1 × 10−15). Pair type was also strongly associated with accuracy (χ2 = 80.1, 14 p < 1 × 10−15), reflecting the substantial reduction in predictive performance observed when moving from population-level to within-sibling comparisons. SNP-set size had a more modest but still statistically significant effect (χ2 = 21.6, p = 2 × 10−5).

In contrast, population structure adjustment method explained relatively little variation in predictive accuracy overall (χ2 = 9.6, p = 0.008), with both the 16-PC and GLMM + 16-PC approaches yielding only slight improvements relative to the unadjusted model. Importantly, the interaction between adjustment method and pair type was not significant (χ2 = 0.008, p = 0.996), indicating that all three methods exhibited comparable levels of attenuation when transitioning from unrelated individuals to sibling pairs.

Together, these results suggest that more sophisticated population structure correction strategies do not substantially improve within-sibling transferability of PRS beyond standard PC adjustment in this dataset, and that reductions in population-level test-statistic inflation do not necessarily translate into gains in causal predictive validity.

Comments (0)

No login
gif