MtlA gene sequencing based on 3rd -generation sequencing for Vibrio typing

Vibrio parahaemolyticus is a species of halophilic gram-negative bacteria found in water, sediment, aerosols, plankton, fish and shellfish. V. parahaemolyticus not only seriously endangers the mariculture industry [1] but can also cause gastrointestinal diseases in humans. Eating food contaminated with V. parahaemolyticus will cause nausea, vomiting, abdominal pain, diarrhea, headache and fever and, in severe cases, may induce septicemia [2], and its global incidence is increasing [3].

Rapid and accurate genotyping of V. parahaemolyticus and understanding its prevalence are highly important for quickly proposing solutions and controlling the spread of diseases and pathogens [4]. For instance, in a study analyzing the genetic and evolutionary relationships of 162 V. parahaemolyticus strains isolated from Guangdong Province, China, researchers utilized multilocus sequence typing (MLST) technology to trace the sources of clinical isolates and prevent the outbreaks of V. parahaemolyticus infection [5]. In another study, researchers performed a genome wide phylogenetic analysis of a global collection of 111 V. parahaemolyticus ST36 isolates and revealed the trend of its transcontinental expansion, which reflects the crucial role of genotyping in monitoring and controlling the spread of pathogens [6]. At present, serotyping and MLST are the main methods for genotyping V. parahaemolyticus.

The serotype of V. parahaemolyticus is detected mainly by mixing the cell antigen O-antigen and the envelope antigen K-antigen with polyvalent antiserum to observe whether agglutination occurs [7]. However, the resolution of the serotype is poor, subjective judgment has a large impact, most strains have the same serotype, or once the corresponding antigen site is mutated, a new serotype may appear, and the real difference between strains cannot be reflected. This series of problems present great challenges for the epidemiological surveillance and tracing of V. parahaemolyticus [8].

MLST is a typing method for amplification analysis of specific gene sequences in the genome. By uploading the sequences of typing genes (generally 6 to 8) to the PubMLST database, the STs of a strain can be obtained through multisequence comparisons with the sequences of existing alleles in the database [9]. The MLST database of V. parahaemolyticus contains seven housekeeping genes, namely, dnaE (557 bp), gyrB (592 bp), recA (729 bp), didS (458 bp), pntA (430 bp), pyrC (493 bp) and tnaA (423 bp). However, for strains with close relatives, it is still difficult to use MLST to achieve accurate traceability [3]. Interspecies recombination events have been reported in the recA gene, leading to undefined MLST types. Furthermore, such recombination in recA can alter the topology of the MLST phylogenetic tree, potentially obscuring true evolutionary relationships [10].

To screen new genes that are easier to standardize, reduce time and cost and can be carried out by grassroots laboratories on their own, we identified mtlA, the core gene of 1,953 bp with the highest variation rate among the genomes of V. parahaemolyticus [4]. This gene encodes PTS mannitol transporter subunit IICBA and can achieve a resolution comparable to or even higher than that of traditional MLST. Furthermore, mtlA shows broad applicability, being present not only in V. parahaemolyticus but also universally detected in 13 pathogenic Vibrio species (e.g., V. cholerae, V. alginolyticus) [11]. Notably, mtlA undergoes positive selection pressure, indicating adaptive evolution associated with environmental survival and pathogenicity. Owing to its economic and computational advantages and universality as a single gene, the mtlA gene may replace MLST as a new gene for typing and tracing with high efficiency.

The MLST scheme was first developed for the species Neisseria meningitides in 1998 [12], and it has been relatively expensive to execute, mainly because of the laborious process of DNA sequencing by Sanger technology at that time [13].

In genotyping, the most central step is gene sequencing. Sanger sequencing, also known as first-generation sequencing technology, has successfully promoted the human genome project, but the shortcomings of low throughput and high cost limit its further large-scale application. Second-generation sequencing technology, also known as next-generation sequencing technology, has been widely used in basic research and clinical diagnosis and treatment because of its advantages of high throughput and low cost. However, the short length of sequencing reads limits their application in genotyping [14].

Third-generation sequencing technology, owing to its advantages of long reading length, fast speed and real-time single-molecule sequencing, has gradually demonstrated clinical application value in tumors, immunity, reproduction and other related fields [15]. Over the past decade, nanopore sequencers, exemplified by the products of Oxford Nanopore Technologies (ONT), have drawn significant attention within sequencing communities due to their abilities to provide long-read sequencing results at a relatively low cost and in a timely manner.

The G-seq500 developed by Geneus Technologies is a new type of nanopore sequencer commercially available recently. Different from the sequencers by ONT that utilizes nanopore as the biosensor to directly read bases of DNA single strand, G-seq500 employs the approach of “sequencing by synthesizing (SBS)”. CW. Fuller et al. and PB Stranges et al. introduced prototype platforms that use nanopores to detect nucleotides one by one during DNA synthesis to determine the sequence of the template DNA [16, 17]. In their work, a DNA polymerase bound with primed DNA template was covalently attached to a nanopore, and the nucleotides were tagged with distinct polymers specific to the A, T, C, G bases. When a nucleotide was captured by the polymerase for synthesis, the tag interacted with the attached nanopore, producing base-specific electronic signal that could be detected by a chip array for real time sequencing. The G-seq500 builds upon a similar strategy, using a nanopore to identify nucleotide bases during DNA synthesis. With extensive optimization of the biochemical system, including a novel nanopore and polymerase, uniquely modified nucleotides, and an improved design of the nanopore-polymerase conjugation, G-seq500 shows significantly enhanced performance compared to the prototype platform, including the read length at the range of tens of kilobase pairs and a remarkably high consensus accuracy exceeding Q50 at the coverage of 50X.

As a third-generation sequencer with superior performance and a viable alternative to products from Pacific Biosciences or ONT, G-seq500 draws our great interest in exploring how its high-accuracy long reads can facilitate genotyping for V. parahaemolyticus. In this study, we tested the efficiency of sequencing the full-length mtlA gene (1,953 bp) via G-seq500, verifying its application in genotyping.

Comments (0)

No login
gif