Figure 1 summarises the approach used in this study for the investigation of the molecular mechanisms of AsymAD based on DEGs, DETs, and DUTs in reference to AD, highlighting the importance of the information obtained from all three methods. Figure 1a schematically represents the brain regions in the datasets used and the pre-processing steps. These regions were selected due to their anatomical proximity and similar functions (e.g. learning, memory, cognition). Figure 1b shows the representative complementary benefits for DEG, DET, and DUT analyses. Some of the transcripts of DEGs may be DUTs, even if they are not DETs, or vice versa. Some genes may be differentially expressed in transcript-level analyses, although they are not DEGs. These cases emphasise the importance of assessing the contribution of each method individually, as well as the value of evaluating the results by combining all three analyses. Figure 1c shows the evaluation steps of the DEG, DET and DUT results through gene list comparisons, enrichment analysis, and subnetwork analysis.
Fig. 1
Overview of the study workflow and the analyses. a The RNA-Seq datasets used in this study and how they were preprocessed to identify AD and AsymAD samples based on their neuropathological and clinical scores. b The first graph shows that Gene A is a DEG while Gene B is not a DEG between the two conditions. The second graph shows DETs based on the three transcripts for Gene A and two transcripts for Gene B. The first transcript of Gene A, T1A, is not differentially expressed, and the remaining transcripts are differentially expressed between the two conditions. Although Gene B is not a DEG, its transcripts are differentially expressed. The third graph shows DUTs according to the relative abundances of transcripts of a gene in both conditions. (eg. transcript usage of T3A is calculated as 40/80 = 0.5 for AsymAD and 80/160 = 0.5 for AD.) The transcript usage of T3A is not changed, while the usage of other transcripts is changed between the two conditions. c The gene lists obtained through DEG, DET and DUT analyses were compared with the AD-related and LMC-related gene lists, subjected to functional enrichment analysis, and used to create a protein subnetwork. * significant expression difference, ns non-significant expression difference, LMC learning-memory-cognition. The brain illustration in this Figure was adapted from OpenClipart-Vectors via Pixabay (https://pixabay.com/vectors/biology-brain-cortex-holiday-1295897/), licensed under the Pixabay Content License
Comprehensive Comparison of DEG, DET, and DUT Analyses for AsymAD Reveals Both Overlapping and Distinct PatternsThe AsymAD groups were not pre-determined in the datasets but were defined in this study according to the clinical and neuropathological scores in both datasets, as explained in Methods, Sect. 2.2. A total of 92 AsymAD individuals for the ROSMAP dataset and 26 AsymAD individuals for the MSBB dataset were identified. In addition, the number of individuals with AD was 211 for ROSMAP and 135 for MSBB based on the consensus between the neuropathological and clinical scores (Fig. 2a).
The differential gene and transcript analyses were evaluated separately for both datasets. The results are shown in Fig. 2b. 995 DEGs were identified in ROSMAP, while there were 320 DEGs in MSBB. There were 1398 DETs in ROSMAP and 2450 DETs in MSBB, while they corresponded to 992 genes (DET-G) and 1979 genes, respectively. For the DUT analysis, there were 293 DUTs in ROSMAP and 541 DUTs in MSBB, and they corresponded to 185 and 425 genes (DUT-G), respectively. For each analysis, the number of common genes identified across the two datasets was 25 for DEG, 128 for DET, and 13 for DUT analyses (Fig. 2c). The genes from these three analyses were combined for each dataset to construct dataset-specific lists, and the similarity of the two datasets was compared. This led to 1631 genes for ROSMAP and 2460 genes for MSBB. The ROSMAP dataset results showed that 15% of its combined gene list overlapped with the genes from the MSBB dataset. The Fisher’s exact test showed that this overlap is highly significant (p = 3.73 × 10⁻⁹) between these two cohorts.
Fig. 2
Comparison of AsymAD and AD groups and differential analysis results in both datasets. a The number of samples in each category for each dataset. b Comparison of DEG, DET and DUT analysis results across ROSMAP and MSBB. c Venn diagrams showing the number of common and different genes between the two datasets for each analysis
To provide a general perspective for the ROSMAP and MSBB results and identify related pathways, combined gene lists from DEG, DET and DUT analyses were analysed in terms of enriched functional terms for each dataset separately (Supplementary Fig. 1 and Fig. 2). Noteworthy findings from the ROSMAP dataset include synaptic processes and plasticity terms like regulation of synaptic plasticity, postsynaptic specialisation, presynaptic membrane, synapse organisation, postsynaptic density and synaptic vesicle cycle; neurotransmitter and synaptic function-related terms like neurotransmitter secretion, GABAergic and glutamatergic synapse, and transmission across chemical synapses; molecular and cellular regulation related terms like regulation of GTPase activity, glial cell differentiation; and cognition. Similarly, enrichment results generated from the gene list created by combining the results from the three analyses in MSBB include synapse-related processes like synapse, glutamatergic synapse, regulation of postsynapse organisation, vesicle-mediated transport in synapse; energy metabolism and mitochondrial function-related terms like ATP binding, mitochondrial inner membrane and lipid metabolic process; and other terms such as microtubule cytoskeleton, cellular response to stress, and heme biosynthesis.
Comparison of the Shared Genes in ROSMAP and MSBB Datasets with LMC-related and AD-related Gene Lists Reveals Strong CandidatesThe list of genes that were found to be significant in at least one of the three analyses in both datasets is shown in Table 1. Since these genes were shared between both datasets, they were considered candidate genes involved in the molecular mechanisms of AsymAD. 230 such genes were identified. The bold genes in the table are highlighted as they are present in the AD-related or LMC-related gene lists. The scoring-based categories of these 230 candidate genes and their distributions to the categories are shown in Fig. 3a. For instance, the R2-M3 category represents genes identified in at least two analyses out of the three analyses performed (DEG, DET, DUT) in ROSMAP and identified in all three analyses in MSBB. The candidate genes were distributed into six categories based on the number of analyses that identified them in each dataset as follows: R1-M1 (153), R1-M2 (11), R2-M1 (55), R2-M2 (8), R2-M3 (1), R3-M1 (2).
Table 1 230 genes that are shared in the ROSMAP and MSBB datasets. These genes were identified by at least one analysis type (DEGs, DETs or DUTs) in both datasets. The letter codes in the category names indicate the datasets, and the numbers show how many analysis types identified the genes for that dataset. The genes in bold represent either AD-related or LMC-related genesThese 230 candidate genes were compared with the LMC-related and AD-related gene lists. The candidate genes include 35 LMC-related and 8 AD-related genes, and two genes are both LMC-related and AD-related. (Fig. 3b). The two genes that were shared by both LMC-related and AD-related genes were NRXN3 (Neuroxin 3) and DGKB (Diacylglycerol Kinase Beta), which are related to synaptic organisation (Bot et al. 2011) and signal transduction (Hozumi and Goto 2012), respectively, highlighting their potential role in neuronal communication, memory formation and neurodegenerative processes.
Some of the 230 candidate genes associated with AD or AsymAD are represented in Fig. 3c, highlighting their involvement in various pathways. Selected pathways in the Sankey plot are GTPase regulator activity, neuroinflammation and glutamatergic signalling, glucagon signalling pathway, long-term potentiation, synaptic membrane, cognition, and learning or memory. The gene with the highest interaction within these pathways is PLCB1 (Phospholipase C Beta 1). This gene is linked to GTPase regulator activity, learning or memory, cognition, long-term potentiation, glucagon signalling pathway, glycerolipid metabolic process, glutamatergic synapse, cholinergic synapse, dopaminergic synapse and neurotransmitter receptors, and postsynaptic signal transmission. The highlighted gene in Fig. 3b, NRXN3, is associated with neuroligin family protein binding, learning or memory, cognition, and the synaptic membrane. The other highlighted gene, DGKB, is less connected and is associated with the synaptic membrane and glycerolipid metabolic process. Other highly interconnected notable genes include CAMK4, PPP3CB, GNG4, RASGRF1, RPS6KA1, and NRXN1 (Fig. 3c).
Mapping 230 candidate genes to the human PPI network revealed that 176 of these genes interacted with one another, forming a subnetwork (Fig. 3d). Genes marked in grey are not part of the original list, but their major function is to interconnect the genes in the list. APP, a gene strongly associated with AD, is one of such genes, and it has a very high number of interactions. In the R3-M1 group, PACSIN2 and S100A6 were identified as strong candidates, significant in three analyses in ROSMAP and one in MSBB. In the R2-M2 group, RIMS1, MRPL1, and CLDN5 were highlighted. RIMS1 was associated with synaptic membrane and GTPase regulator activity (Fig. 3c). In the R2-M1 group, CAMK4, PLCB1, COL26A1, NRXN3, RASGRF1, MRPL30, FGF13, and GFAP were highlighted. CAMK4 is associated with pathways such as learning or memory, cognition, long-term potentiation, activation of NMDA (N-methyl-D-aspartate) receptors and postsynaptic events, neurotrophin signalling pathway, cholinergic synapse and neurotransmitter receptors, and postsynaptic signal transmission. In the R1-M2 group, GNG4 and FBL genes were featured in the subnetwork. GNG4 was associated with pathways including glutamatergic synapse, cholinergic synapse, GABAergic synapse, dopaminergic synapse, neurotransmitter receptors and postsynaptic signal transmission (Fig. 3c). The genes highlighted in the R1-M1 group outside this category are COL25A1, PCOLCE, RPL19, and ENPP5. A subset, as a circle within the R1-M1 group, represents genes that were not identified in the DEG analyses but were significant in the transcript analyses (DET and DUT) and also part of the LMC-related or AD-related gene lists. Those genes were obtained by first identifying genes consistently determined as DETs in both datasets and genes consistently determined as DUTs in both datasets. Then, genes that were in the DEG lists of either dataset were excluded from the common DET and common DUT genes and intersected with the LMC and AD-related genes (Supplementary Fig. 3). The LMC or AD-related genes in this small circle are GABRA1, NR1H2, ITSN1, NME1, PFKP, RAB3GAP1, PDXK, FGF1, NAPA, NDRG2, IQSEC2, PPP3CB, and IDH3A. These genes prove that transcriptomic studies relying solely on DEG analysis can miss transcript-level changes.
Fig. 3
Generation and evaluation of 230 candidate genes. a Genes derived from the DEG, DET, and DUT analyses were evaluated across nine combinations for the two datasets. b A Venn diagram shows the comparison of 230 candidate genes with the AD-related and LMC-related gene lists. c Sankey plot showing specifically selected pathways and the candidate genes related to these pathways. d The subnetwork results of the 230 candidate genes mapped onto a human PPI network. 176 of them mapped and exhibited strong interactions with each other. Purple nodes represent the R1-M1 group, while the smaller purple circle indicates genes originating exclusively from the DET or DUT analyses, but not from the DEG analysis, and are also LMC- and AD-related. Blue nodes represent the R1-M2 group, green nodes represent the R2-M1 group, yellow nodes represent the R2-M2 group, and pink nodes represent the R3-M1 group. The genes indicated in grey are the genes not within the list of the candidate genes but included due to the K parameter of the subnetwork discovery algorithm, since they connect the candidate genes
Additionally, only the intersections of the transcripts, not the genes, in the DET and DUT analyses were compared across the two datasets to reveal the differential transcripts that were compatible across the two datasets (Supplementary Fig. 4). It was found that eleven transcripts were common in the DET analysis, and four transcripts were common in the DUT analysis. The first example from these genes is the transcript ENST00000366597 of GNG4, which is significantly changed in only DUT analysis for ROSMAP and significantly changed in DET and DUT for MSBB (Fig. 4). The second gene is the transcript ENST00000274609 of ADAMTS2, which is significantly changed in DEG and DET analyses for ROSMAP and significantly changed in DEG, DET, and DUT analyses for MSBB (Supplementary Fig. 5). Moreover, in DET, one of the transcripts of the PCOLCE gene (ENST00000472348) was downregulated in AsymAD (Supplementary Fig. 6a), while in DUT, the proportion of a transcript of the ENPP5 gene (ENST00000230565) is higher in AsymAD (Supplementary Fig. 6b) in both datasets. The MRPL1 gene, which is one of the genes that are not associated with the AD-related and LMC-related gene lists, is not significant in the DEG analysis but has significantly changed transcripts: in the ROSMAP dataset, ENST00000315567 (DET) and ENST00000504901 (DUT); and in the MSBB dataset, ENST00000515625 (DET and DUT) (Fig. 5). This gene is a prominent example that proves how conventional DEG analysis can miss transcript-level information.
Fig. 4
DEG, DET, and DUT analysis results for GNG4. Transcript expressions and transcript usages of the GNG4 gene are significantly different between the AsymAD and AD groups in both datasets. The empty dot plots represent the transcripts of the GNG4 gene that are not expressed or lack transcript usage estimates. * Significantly Expressed, NS: Not Significant, NE: Not Expressed, NA: Not Available
Fig. 5
DEG, DET, and DUT analysis results for MRPL1. Transcript expressions and transcript usages of the MRPL1 gene are significantly different between the AsymAD and AD groups in both datasets. * Significantly Expressed, NS: Not Significant, NE: Not Expressed, NA: Not Available
The results obtained from the Control-AsymAD and AsymAD-AD comparisons for ROSMAP and MSBB datasets are presented in Supplementary Table 4. To better define the molecular characteristics of the AsymAD category, genes/transcripts were categorised into three groups: those altered only between Control-AsymAD (remaining unchanged between AsymAD-AD), those altered only between AsymAD-AD (remaining unchanged between Control-AsymAD), and those showing differences in both comparisons. Within these categories, all significantly altered genes/transcripts in DEG, DET or DUT analyses are listed in Supplementary Information Files 3 and 4 along with their respective cohort information. The distribution of these categories is shown in Supplementary Fig. 7.
These results demonstrate that the majority of the identified genes/transcripts originated from the AsymAD-AD comparison, supporting the main focus of our study. The comparisons allowed the following interpretations: (i) genes altered only between Control-AsymAD likely represent early pathological changes independent of clinical AD symptoms; (ii) genes altered only between AsymAD-AD may represent mechanisms related to resilience during the transition to AD; and (iii) genes altered in both comparisons likely reflect the genes specific to the AsymAD mechanism (Supplementary Information File 5). Subsequently, the 230 candidate genes identified in the AsymAD-AD comparison were cross-referenced to determine if they were altered specifically in the AsymAD-AD comparison, in the Control-AsymAD comparison, or both (Supplementary Information File 6). MRPL1, ADAMTS2, GNG4, PCOLCE, and ENPP5 genes, shown within the boxplot figures (Figs. 4 and 5 and Supplementary Figs. 5, 6), were observed to be differentially changed only in the AsymAD-AD comparison (Supplementary Table 5).
Cross-Species Functional Context of Candidate Genes Using a Mouse DatasetIn order to further validate our candidate genes, a transcriptome dataset obtained from mouse models representing AD and AsymAD was analysed in this section (Fig. 6a). In the dataset (GSE147434) (Pérez-González et al. 2020), 696 DEGs were obtained by comparing gene expressions of memory-impaired and memory-intact mice with AD pathology. The 696 genes were enriched in signalling, response to stress, apoptotic process, and most importantly, sensory perception, which is consistent with behavioural assessments such as the Morris Water Maze and Fear Conditioning tests used in the study. We checked the relationship between 230 candidate genes identified in the previous section and the differential genes from the mouse dataset. In the 230 gene list, 45 genes intersected with AD-related or LMC-related genes (See Fig. 3b). However, the remaining 185 genes may also play a role in progressing from AsymAD to AD, offering novel molecular insights into AsymAD. Therefore, we focused on the list of 185 genes for validation.
Among the 185 genes that were not AD-related or LMC-related, four (SLC6A12, UXS1, APOL3 and SPATA13) matched with the 696 DEGs from the mouse dataset. These genes may serve as novel biomarker genes for AsymAD. We next evaluated the 185 genes and 696 mouse DEGs in terms of their regulatory elements and biological pathways in order to identify potential shared mechanisms. Although only a small number of the 185 candidate genes overlapped with the mouse DEGs, we showed that these two gene sets are regulated by the same transcription factors (TFs), linked to the same microRNAs, and involved in the same biological pathways (Fig. 6b). The Jaccard similarity analysis shows that the TFs of the two gene sets are 90% similar, miRNAs are 60% similar, and related pathways are 43% similar (Fig. 6b). Moreover, Fisher’s exact tests confirmed that these overlaps in all three categories are highly significant (p < 2.2 × 10⁻¹⁶). We further checked if there were any known physical interactions between the proteins of the 185 genes and the DEGs from the mouse dataset. As a result, 71 of the 185 genes had physical interactions at the protein level with 141 of the mouse DEGs (Supplementary Table 6, Fig. 6c). Our analysis shows that even if the overlap between these gene sets is limited, the majority of the 185 genes either have direct physical interactions with the mouse DEGs share the same TF or miRNA regulators, or are involved in the same pathways.
Fig. 6
Validation with the mouse dataset for the identified candidate genes. a A workflow that shows each step applied to validate candidate genes. b Commonly regulated TFs, associated miRNAs, and involved pathways are shown between the mouse DEGs and 185 candidate genes. c The interaction network that shows the relation between the mouse DEGs and candidate genes. The outer circle represents mouse genes, and the inner circle represents the candidate genes. The mouse illustration in Fig. 6a was adapted from Clker-Free-Vector-Images via Pixabay (https://pixabay.com/users/clker-free-vector-images-3736/), the GEO database logo is reproduced from NCBI GEO (https://www.ncbi.nlm.nih.gov/geo/), and the icons and template in Fig. 6a were created using Canva Pro (https://www.canva.com)
Comments (0)