A detailed summary of the In-silico analysis results is provided in Supplementary Tables 1–3 as well as supplementary files 1 and 2.
No common variants with a MAF > 0.01 (1000G global MAF) were found in the PCR primer binding regions. However, it needs to be taken into consideration that this filter in the NCBI variation viewer refers to the global allele frequency. Thus, among the many SNPs with a global MAF < 0.01 identified in the primer binding regions (Supplementary File 2), some may be (more) common in certain populations (cf. Table 4).
Table 4 Detected SNPs in primer binding regionsFor the SNaPshot™-based analysis method, it should be noted that the amplicons generated by the assay 01 primer sets 2 and 3 contain 3 and 6 SNPs respectively. There are thus a few SBE primers overlapping each other, for which other HIrisPlex-S-SNPs can be found within the primer binding regions (Supplementary Tables 2, 3), and which may thus impact each other’s genotyping results.
To address this issue, ‘wobble’ bases [34] could be introduced during SBE primer design, ensuring that a proportion of the SBE primers could efficiently bind to the target region independent of the SNP allele carried by the individual. However, no participating laboratory using the SNaPshot™ method describe having their SBE in this way.
The majority of PCR primers do not pose a high risk of generating additional off-target PCR products (Supplementary Table 1). There are, however, three primer pairs located within the two genes HERC2 and TYR, for which off-target PCR product generation was considered “high-risk”:
The assay 02 PCR primer sets 6 and 8 (amplifying regions around rs6497292 and rs1667394) are intended to amplify the genomic locations chr15: 28,250,982–28,251,131 and chr15: 28,285,002–28,285,131, respectively, within the HECT And RLD Domain Containing E3 Ubiquitin Protein Ligase 2 (HERC2) gene. With a single mismatch for one primer of set 6 and one mismatch in each primer of set 8, the primer sets 6 and 8 potentially amplify six and three different regions on chromosome 15, respectively. In each case, this could lead to the generation of an off-target PCR product of identical length and high sequence similarity compared to the intended PCR product (Supplementary File 1). The high sequence similarity is explained by the fact that several partially duplicated paralogs of the HERC2 gene can be found on chromosome 15 [35, 36]. Stringent PCR conditions are thus important to avoid/reduce the off-target amplification. If nevertheless co-amplified, it would be possible to differentiate between the target and off-target PCR products in MPS-based genotyping as the PCR product sequences are not entirely identical. Within the HPS-MPS-MiSeq panel, the reverse PCR primer for the rs1667394 target was modified and now contains two mismatches to the off-target regions, thus reducing the risk of co-amplification.
For the SNaPshot™ assay, the probability of generating an off-target SBE product may also be considered low, as the rs6497292-SBE-primer can only bind to the respective off-target PCR products with at least two mismatches in the primer-binding region and the rs1667394-SBE-primer can only bind to the respective off-target PCR products with at least three mismatches, thus off-target SBE products will be generated (much) less efficiently than the intended SBE products.
The assay 02 PCR primer set 9 (amplifying a region around rs1126809) is intended to amplify the genomic location chr11: 89,284,739–89,284,838, lying within the tyrosinase gene (TYR). With a single mismatch for each primer, these primers potentially amplify a second region on chromosome 11 (49415689–49415788). This would lead to the generation of an off-target PCR product of identical length and high sequence similarity compared to the intended PCR product (Fig. 1 and Suppl. File 1). The high sequence similarity is explained by the fact that the off-target genomic region lies within the tyrosinase like pseudogene (TYRL) [37]. As the PCR product sequences are not entirely identical, it should be possible to differentiate between the target and off-target PCR products by MPS-based genotyping methods if not only the SNP of interest is considered during interpretation. However, as the rs1126809-SBE-primer can bind to the off-target PCR product without a single mismatch, the genotyping results obtained from SNaPshot™-based methods may be impacted by the off-target binding of the PCR primers.
Fig. 1
Determination of cause for wrong additional G detection in rs1126809 (TYR) (A) Schematic overview of experimental approach (B) Sequence of region of interest for TrACE 2024 extended individual 1 with primer pair F/R2individual B (upper pane) and F/R1 lower pane). Orange: SNP, black squares: additional discriminant positions between TYR and TYRL. (C) Sequence comparison of target (TYR) and off-target region containing TYRL (highlighted in grey). Orange/Black squares: positions as in chromatogram in B
Stringent PCR conditions can minimize the risk of co-amplification of off-target regions in these cases. Nevertheless, care must be taken as spurious amplification might be possible. Normally, extreme imbalances in allele coverage (MPS) or peak height (SNaPshot™) could be an indication for unintended off-target amplification. However, the interpretation may become more difficult in cases of low-quality/low-quantity samples which may exhibit increased allelic imbalances.
Discrepant results for samples in FDP collaborative exercisesDiscrepant genotyping results were observed for individual SNPs in the FDP modules of the GEDNAP (2022) and TrACE (2023, 2024) collaborative exercises. As these discrepancies were not only reported by a single or a few labs, but by a larger number of participating laboratories, a systematic investigation was performed.
Discrepant genotyping results obtained during the three collaborative exercises as well as a summary of the analysis methods employed by the respective participating laboratories are provided in Table 2.
For the GEDNAP 65 collaborative exercise, less detailed information was available, as participating laboratories were not yet required to provide raw data (electropherogram or MPS retrieved SNP information) at that time. Due to the observed discrepancies, this was changed for subsequent collaborative exercises (especially the TrACE exercise, first conducted in 2023).
Discrepant results for the SNP rs1126809 were obtained for three different individuals from two independent collaborative exercises. In-silico analysis had shown that the original HIrisPlex-S PCR primers amplifying the region around this locus can potentially amplify a second PCR product with one mismatch in each primer pair (Fig. 1, Suppl. File 1, Suppl. Tables 1, 2). This off-target PCR product contains a G nucleotide at the sequence position corresponding to rs1126809 within the intended PCR product. This G nucleotide will only lead to wrong genotyping results for individuals carrying a homozygous A allele at the true rs1126809 genomic location. It is thus plausible to assume that the incorrect reporting of the AG genotype by some laboratories was caused by unintended amplification of a genomic region within the TYRL pseudogene. Indeed, we were able to confirm this assumption by targeted sequencing analysis with newly designed primers (Table 3; Fig. 1): When the forward primer was used in combination with reverse 2, a second off-target PCR product was co-amplified, distinguishable from intended PCR product by sequence analysis. However, when the forward primer was used in combination with reverse 1, which carried two mismatches to the off-target-region, only the intended PCR product was generated, allowing for a verification of the AA as the correct genotype at rs1126809 (Fig. 1) for the GEDNAP 65 A and C individuals as well as TrACE 2024 ext. ind. 1.
All MPS-performing laboratories obtained the correct results for rs1126809 (TYR) during the GEDNAP65 as well as TrACE 2024 (Table 2). It is not clear, if background co-amplification nevertheless occurred but remained below the reporting threshold, if the primers (off the commercial) kits are modified avoiding any co-amplification, or if the PCR conditions were strict enough to avoid any remaining annealing to the off-target region. Deeper analysis was possible for our own MPS analysis, which was based on the HPS-MPS-MiSeq panel, therefore using the original HIrisPlex-S primer. In all three affected samples (GEDNAP 65 individual A and C, and Trace 2024 ext. individual 1, the presence of the G nucleotide was noticeable. However, using the HPS-MPS analysis pipeline from Breslin et al. [18], only the AA genotype was called as final genotype. This is due to the strong allele imbalance of 88–95% A vs. 12 − 5% G in all three cases (e.g. final coverage after HPS-MPS analysis pipeline application for GEDNAP 65 individual C in duplicate analysis: A: 3483/1449, G: 211/46). Furthermore, MPS allowed a deeper investigation of the amplicon sequence itself, which led to the detection of multiple differences in the amplicon sequence compared to the TYR reference (cf. Figures 1 and 2), supporting the final decision to not report the G allele. Exemplarily, the IGV is presented for the ext. individual 1 of the TrACE 2024 exercise (Fig. 2).
Fig. 2
MPS (left) and SNaPshot (right) genotyping results for rs1126809 (TYR) For sample GEDNAP 65 individual C, the co-amplified G nucleotide can be connected in MPS to the off-target pseudogene TYRL which also has a T nucleotide at position chr11:89.017.973 (in contrast to the C in TYR) (highlighted in red box). In the samples GEDNAP 64 individual C and GEDNAP 65 individual B, some reads of the off-target TYRL can also be seen in the MPS data (as can by differentiated based on the T at chr11:89.017.973), which does however not impact genotyping results at position chr11:89.017.960 as the genotype of TYR also contains a G for these samples. The SNaPshot™ electropherogram also shows the additional G nucleotide of TYRL in GEDNAP65 individual C (It may be assumed that a proportion of the G signal for GEDNAP 64 individual C and GEDNAP 65 individual B also originates from the co-amplified TYRL, however, this cannot be differentiated in the SNaPshot™ data)
Laboratories using the SNaPshot™-based genotyping methods obtained more divergent results in both collaborative exercises (Table 2). Again, the use of different amplification conditions might have led to reduced co-amplification of the off-target PCR product. Details regarding the PCR chemistry and employed amplification conditions were not obtained by the organisers of the collaborative exercises. Additionally, laboratories might have used different strategies (analytical thresholds, thresholds for heterozygote imbalance), which might have led to discrepant assessment of the co-amplified G-nucleotide. Indeed, in our own SNaPshot™ analysis (including the original HIrisPlex primer for the analysis of rs1126809), small G-nucleotide-peaks next to larger A-nucleotide peaks (71–91% A vs. 9–29% G in all three cases, e.g. for GEDNAP 65 individual C peak heights: A = 2545 rfu, G = 238 rfu), which led us to report the genotype as AA (Fig. 2).
Discrepant results for the SNPs rs10756819 and rs1470608 were obtained for a single individual in one collaborative exercise, each. In-silico analysis did not suggest relevant off-target primer binding sites. No common variants (MAF ≥ 0.01) were discovered in the respective PCR or SBE primer binding locations. However, Sanger sequencing of the regions of interest revealed heterozygous primer binding site mutations within the HIrisPlex-S PCR primers for both rs10756819 (within the gene basonuclin-2 (BNC2), (Fig. 3)) and rs1470608 (within the gene oculocutaneous albinism type 2 (OCA2)) (Table 4).
Fig. 3
Determination of a primer binding mutation in the HIrisPlex-S primer region for rs10756819 detection (BNC2) (A) Schematic overview of experimental approach (B) Sequence of region of interest for GEDNAP65 individual B (upper chromatogram). In addition, the sequence of GEDNAP64 individual B (with the same AG genotype in the BNC2 SNP rs1075819) without a primer binding mutation is shown (lower chromatogram)
It was observed that all participants using the SNaPshot™-based genotyping method reported the incorrect homozygous genotyping results for rs10756819 and rs1470608, respectively. Although this information was not provided in detail, it can be assumed that all laboratories used either the original HIrisPlex-S primer panel or a custom assay based on these primers.
In case of MPS, nearly all laboratories obtained concordant results for all three samples. Exceptions are two laboratories which were not able to detect the heterozygote genotype for rs10756819 and rs1470608. One of these two was our own laboratory. Here, the HPS-MPS-MiSeq was used, which contains the same primers at the SNPs affected primer binding positions resulting in non-detection of the heterozygous alleles. Even after an in-depth sequence analysis, no background amplification of the second allele was detectable in our data. It is not known on which MPS-assay the second discordant result in BNC2 is based. The second discordant result in OCA2 was obtained in case of the use of the ForenSeq DNA Signature Prep Kit (Verogen). However, correct results were observed for the ForenSeq Imagen kit containing the same primers (personal communication from Verogen). This might be due to different calling settings or differences in the PCR conditions between laboratories.
Common SNPs in PCR primer binding regions should be avoided during primer design. However, avoiding the occurrence even of any rare SNP in a primer binding region is hardly possible. The possibility of such occurrences is more difficult to assess for assays for which the primer sequences are not publicly available.
Impact of discrepant genotyping results on phenotype predictionsLastly, we aimed to assess the impact of incorrect genotyping results on the obtained phenotype predictions. The three SNPs for which discrepant genotyping results were obtained during the collaborative exercises were all part of the skin colour prediction model only, thus eye and hair colour predictions were not affected. The visual phenotypes and the predicted probabilities obtained with the HIrisPlex-S-model [16] for each of the five skin colour categories depending on the genotype input are displayed in Table 5. In most cases, the incorrect genotyping result had a minor impact on phenotyping results as it did not lead to a different verbalisation of the predicted phenotype as suggested by the HIrisPlex-S-manual provided with the prediction webtool [16]. In two cases, a slightly larger difference was observed: For individual B from GEDNAP 65, a “dark to dark-black” skin colour would be predicted with the correct GA genotype, whereas the incorrect GG genotype would lead to the prediction of a “dark-black” phenotype. For individual C from GEDNAP 65, a “pale to intermediate” skin colour would be predicted with the correct AA genotype, whereas the incorrect AG genotype would lead to the prediction of a “pale” phenotype. However, it should be kept in mind that the predicted phenotype depends on the combination of all alleles detected for the respective individual, thus, the impact of individual genotyping errors on phenotyping results may be larger in other cases.
Table 5 Predicted probabilities for individuals with discrepant genotyping results for each of the five skin type categories obtained with the Hirisplex-S-webtool [16]Overall, it should be kept in mind that the majority of samples used to train the HIrisPlex-S phenotype prediction model [16] were genotyped using the SNaPshot-based method on capillary electrophoresis with the original HIrisplex-S primer sets. It thus remains questionable whether the use of methods allowing to obtain more correct genotyping results (e.g. by being able to unambiguously differentiate between amplicons generated from the TYR gene and the TYRL pseudogene with MPS-based methods or being able to mitigate the impact of primer binding region mutations by using adapted primer sequences) generally increases phenotype prediction accuracy. This effect might be more pronounced for variants more common in non-European populations, which are underrepresented in the HIrisPlex-S training dataset and for which a genotyping error in the training dataset can thus be expected to have a more severe impact on model prediction performance.
To be able to assess the impact of potential errors in the HIrisplex-S training data set (arising from genotyping biases due to primer binding site mutations, the overlap of SBE primers or off-target amplification of the TYRL pseudogene), the MPS-based verification and/or extension of the training dataset would be desirable.
In case of ambiguous or potentially divergent results, obtaining phenotype predictions with different input options for the affected SNP to assess the impact on phenotyping results is recommended. However, for users of community or commercial primer panels, it is not always possible to assess whether and how these differ from the original HIrisPlex-S primer panel and thus might potentially produce divergent genotyping results (e.g. in case of primer binding site mutations). For the sake of genotyping consistency, we thus highly recommend that custom or commercially designed primer sequences should be made publicly available or at least that deviations from primer sequences used to generate the training data set should clearly be communicated.
Comments (0)