A total of 115 samples from 10 couples who underwent routine PGT-M for α-thalassemia at Nanning Second People’s Hospital from 2021 to 2023, along with their corresponding embryos were enrolled in this retrospective study. These 10 couples carried diverse SVs and SNVs including –SEA deletion, -α3.7 deletions, -α4.2 deletions, and αcsα (Table S1). Genomic DNA was extracted from their peripheral blood with the QIAamp DNA blood mini kit (Qiagen, Hilden, Germany). For embryos, MDA was performed with several trophectoderm cells with the REPLI-g® Single Cell Kit (Qiagen, Hilden, Germany).
Control methodThe genotypes of all these couples and their corresponding embryos have been identified by control methods (Fig. 2). Briefly, genomic DNA from the parents and their relatives was used to confirm variant mutation by Sanger sequencing and Gap-PCR, and haplotypes were constructed by NGS. Then, MDA products of embryos were screened for aneuploidy through NGS using the MGISEQ-2000 platform (MGI Tech Co. Ltd., Shenzhen, China) at approximately × 0.1 genome depth. Next, linkage analysis was conducted on MDA products of embryos by NGS, and mutation was validated by Sanger sequencing and Gap-PCR. Finally, the genotype of embryo was determined.
Fig. 2
Flowchart and comparison between tlrPGT-α-thal and control methods. PGT, preimplantation genetic testing. gDNA, genomic DNA. MDA, multiple displacement amplification. NGS, next-generation sequencing
Design of the LRS-based PGT for α-thalassemia (tlrPGT-α-thal)We aimed to establish a PGT-M method for α-thalassemia, which consists of haplotype construction, linkage analysis, and genotype determination of the embryo without requiring a proband. Two sets of primers were designed in order to avoid the influence of ADO (Fig. 1A). The first set of primers was used to detect SVs and SNVs involved in HBA genes, named, HBA-F, HBA-R, S-F, STF-R, and TF-F, which was modified from our previous approach (CATSA) [23] to adapt for PGT-M applications. The second set consisted of four pairs of primers (L1-F and L1-R; L2-F and L2-R; L3-F and L3-R; L4-F and L4-R) targeting regions within 1 Mb upstream and downstream of the HBA core region to analyze the haplotypes of the peripheral region.
The selection of the second primer set was based on population allele frequencies of SNPs and their amplification efficiency. First, the population allele frequencies of all SNPs located within the 1 Mb region of the upstream and downstream of HBA gene were listed according to gnomAD (https://gnomad.broadinstitute.org/). Secondly, SNVs with appropriate population allele frequency (40–60%) were considered to be optimal. Thirdly, 6–10 kb genomic areas with maximum number of SNPs were selected as amplicon regions. Fourthly, primers were designed to avoid regions containing heterozygous SNPs to avoid allele-specific amplification failure caused by primer-template mismatches. Fifthly, the selected genomic regions were prioritized, and the primers compatible with the entire system were subsequently chosen through the amplification validation. Finally, the four optimal regions were selected based on both assay performance and cost-effectiveness.
By analyzing the haplotypes of parents and embryos in the five regions above, the haplotype linkage between the HBA and the peripheral regions was established to determine the genotypes of all embryos.
Long-range PCR and PacBio sequencingA minimum of 10 ng of genomic DNA and 100 ng of MDA products (Fig. 1B) were co-amplified in a single multiplex long-range PCR (LR-PCR) reaction using both primer sets (mentioned above), following a two-step amplification method (Fig. 1C). Multiplex LR-PCRs were performed in 50-μL reactions containing 10 to 100 ng of genomic DNA or MDA products, 1 × PCR KOD FX Neo buffer, 0.4 mM of each dNTP, 1 μM of primer mixture, and 1 μL of KOD FX Neo (Toyobo, Osaka, Japan). PCR program for optimal fragment amplification was 94 °C for 2 min (1 cycle), 98 °C for 10 s, and 68 °C for 14 min (30 cycles) and 68 °C for 10 min (1 cycle) (Fig. 1D).
The single-molecule real-time (SMRT) dumbbell (SMRTbell) libraries were prepared and sequenced as previously described [23] (Fig. 1E,F). Briefly, the targeted amplicons (about 200 ng per sample) were subjected to a one-step end-repair, phosphorylation and ligation to ligate barcoded adaptors using T4 DNA polymerase, T4 polynucleotide kinase, and T4 DNA ligase. The fragments failed to ligation were digested by exonucleases to reserve only the dumbbell-shaped fragments. The individual pre-libraries of each sample were purified by PB beads (Pacific Biosciences, Menlo Park, CA, USA) and pooled together by equal mass. The mixed library was combined with sequencing primer and Sequel® II Polymerase 2.0 (Pacific Biosciences, Menlo Park, CA, USA) to obtain a sequencing library. The polymerase-bound SMRTbell complexes were sequenced on the Sequel® II LRS system under the circular consensus-sequencing (CCS) mode for 30 h.
PacBio date analysis and the haplotype linkage analysisCCS reads were split by barcoded adaptors for different samples and aligned to reference genome hg38with pbmm2 (https://github.com/PacificBiosciences/pbmm2) (Fig. 1G). The quality control criterion required a minimum sequencing depth of > 100 × for aligned reads. FreeBayes1.3.4 (https://www.geneious.com/plugins/freebayes; Biomatters, Inc., San Diego, CA, USA) was used to call SNVs and indels. SVs of HBA were called based on HbVar databases (https://globin.bx.psu.edu/hbvar/menu.html), Ithanet databases (https://www.ithanet.eu/), and LOVD databases (https://www.lovd.nl/).
In addition to the haplotype calling of HBA gene, the linkage relationship analysis was also used to determine the genotypes of embryos (Fig. 1H). First, haplotypes for all samples in the five regions were determined using whatshap (https://github.com/whatshap/whatshap), based on the informative heterozygous SNPs within each fragment. Then, by comparing haplotypes between parents and embryos in the four peripheral regions, a linked haplotype profile spanning the four regions was established. In order to avoid the impact of ADO, a retrieval-based method was developed to search the haplotypes with specific SNP markers from the father or mother in all embryos of one family in the HBA region, and the genotypes of the selected embryos can link the haplotypes of HBA to the four peripheral regions. Finally, the genotypes of the remaining embryos were determined based on the haplotype linkage relationship.
Sequences were visualized by Integrative Genomics Viewer (IGV) (Fig. 1I).
Comments (0)