Lianming Du, Lian Wang, Jiahao Chen, Songwen Tan, Peng Guo*, Qin Liu*. Reference genome assembly of the mandarin rat snake (Euprepiophis mandarinus). Journal of Heredity, 2026, 117(4): 834–842.
The mandarin rat snake (Euprepiophis mandarinus), a widely distributed and ecologically important species endemic to Asia, serves as an ideal candidate for establishing a reference genome for the genus Euprepiophis. This is due to its wide distribution, and its overlapping distribution with and morphological similarities to the endangered Euprepiophis perlaceus. Here we present the first high-quality, chromosome-level reference genome of E. mandarinus, assembled using PacBio HiFi long-read and Hi-C sequencing. The final assembly has a length of 1.65 Gb contained in 56 scaffolds, a scaffold N50 of 210 Mb, a benchmarking universal single-copy orthologs completeness of 98.2%, and 99.12% bases anchored onto 19 chromosomes. We identified 22,332 protein-coding genes, 3,440 noncoding RNAs, and approximately 7 million repeat elements accounting for 58.38% of the assembled genome. This reference genome provides a valuable resource for taxonomic, phylogenetic, comparative, and conservation studies of the genus Euprepiophis.
Jiahao Chen, Qin Liu, Songwen Tan, Peng Guo*, Lianming Du*. De Novo Assembly and Characterization of Venom Gland Transcriptome for Rhabdophis lateralis. Toxins, 2026, 18(4): 167.
Rhabdophis lateralis is a snake species within the family Natricidae, which is widely distributed across mainland China, Russia, and Korea. Although this species was once thought to be non-venomous, there are quite a few cases demonstrating its bite could be fatal. In this study, we performed de novo assembly and analysis of the transcriptome data from the Duvernoy’s gland of R. lateralis, aiming to characterize its venom transcriptome and reveal the molecular basis of its toxicity. Among 6196 annotated transcripts, 77 were identified as potential toxin transcripts belonging to 26 toxin families. The most highly expressed toxin family was the SVMP family, accounting for 51.10% of the total toxin expression. The other notable toxins included cysteine-rich secretory proteins (CRISPs, 22.36%), c-type lectins (CTLs and snaclecs, 12.13%), and three-finger toxins (3Ftxs, 6.36%). Phylogenetic analyses indicated that SVMPs, CRISPs, and three-finger toxins (3FTxs) are evolutionarily conserved within Colubridae, whereas CTLs likely arose through convergent evolution. All identified SVMPs were classified as P-III type, with one sequence displaying a unique deletion distinct from conventional truncation patterns. The predominantly expressed CTLs are more likely to combine into dimers, exerting coagulation activity. This study provides an insight into the toxin gene expression in the Duvernoy’s gland of R. lateralis, which will benefit future research into the ecological and pharmacological significance of toxins in the genus Rhabdophis.
Mengna Feng†, Xiaoyu Wu†, Xin Hu, Yi Wu, Shiyi Gou, Qiman Ran, Yang Yuan, Ting Huang, Lufeng Dan, Yiwen Chu, Xikun Zhou, Kelei Zhao*, Lianming Du*. Repurposing Tirazone as an effective quorum-sensing inhibitor against Pseudomonas aeruginosa virulence and biofilm formation. Journal of Antibiotics, 2026, 79(4): 248–263.
Antibiotic resistance has emerged as a critical global public health challenge. Quorum sensing (QS), a density-dependent regulatory mechanism, plays a pivotal role in bacterial pathogenesis by coordinating virulence factor expression, making it a critical target for antivirulence therapy. Leveraging a drug repositioning strategy, this study investigated the antivirulence potential of drugs in the database of DrugBank on the common opportunistic pathogen Pseudomonas aeruginosa by virtual screening. Molecular docking analysis predicted that the antitumor drug, Tirazone, could bind to the core QS regulatory proteins, LasR, RhlR, and PqsR of P. aeruginosa with abundant active sites, whereas the binding free energies were higher than those of the native QS signals. In vitro experiments demonstrated that Tirazone significantly suppressed virulence factor secretion, cell motilities, and biofilm formation in the model P. aeruginosa strain PAO1, and downregulated the expression of a series of QS-related genes with low effective concentration (≤ 8 μM). A competitive binding model of QS signal molecules further elucidated that Tirazone interfered with QS signaling by competitively inhibiting the function of LasR, RhlR, and PqsR. Additionally, Tirazone treatment significantly protected Caenorhabditis elegans and mouse models from P. aeruginosa infection, and reduced the bacterial loads and pathological lesions in mouse lungs. Moreover, Tirazone demonstrated synergistic effects with polymyxin B, levofloxacin, and amikacin, significantly enhancing their bactericidal efficacy in treating P. aeruginosa. This study reveals the molecular mechanism underlying Tirazone’s multi-target intervention in the QS system, and provides an experimental foundation for developing combination therapies based on antivirulence strategies.
Qingqing Zhang†, Cheng Chen†, Xiaoqin Mu, Zihan Zhang, Cuiling Luo, Chenjuan Zeng, Bisong Yue, Zhenxin Fan*, Lianming Du*. Alleviation of Ulcerative Colitis in Mice by Individual Fermentation of Periplaneta americana Powder with L. bulgaricus SN22 and S. thermophilus SN05. Microorganisms, 2026, 14(2): 301.
The escalating global incidence of ulcerative colitis (UC) underscores the demand for novel therapeutic strategies. This study investigated the fermentation of Periplaneta americana (PA) powder using two conventional dairy starter strains, Lactobacillus delbrueckii subsp. bulgaricus SN22 and Streptococcus thermophilus SN05, to enhance its functional properties, particularly anti-inflammatory activity, via microbial processing. Both strains demonstrated favourable safety and antimicrobial activity. Untargeted metabolomics revealed that fermentation significantly altered the metabolite profile of the PA supernatant, enriching compounds with potential bioactivities, notably anti-inflammatory (e.g., 3-anisic acid) and antioxidant (e.g., vitamin U) properties. In the DSS-induced mouse colitis model, treatment with the fermented supernatant alleviated intestinal inflammation compared to the unfermented group. This was demonstrated by significantly reduced levels of the pro-inflammatory cytokines IL-1β and TNF-α, along with improved maintenance of intestinal barrier integrity. Further in vitro assays showed that the fermented supernatant significantly suppressed proliferation and clonogenicity in human HT-29 colon cancer cells, while also inducing reactive oxygen species accumulation and apoptosis. Results demonstrate these strains are multifunctional starters possessing superior antimicrobial and anti-inflammatory efficacy. This study employed LAB fermentation of insect-derived matrices to derive bioactive components. The fermentation products exhibited anti-inflammatory potential, offering a potential microbial transformation strategy for developing functional products for adjunctive UC intervention.
Jiahao Chen, Qin Liu, Songwen Tan, Peng Guo*, Lianming Du*. A chromosome-level genome assembly and annotation for the beauty snake Elaphe taeniura. Journal of Heredity, 2026, 117(1): 141-150.
The genus Elaphe Fitzinger, a large species clade within Colubridae, comprises 18 non-venomous snake species. Among them, Elaphe taeniura serves as a representative species due to its wide distribution and strong adaptability. However, genomic studies on this group remain limited. Here, we present a chromosome-level, high-quality reference genome for this species. The genome is assembled by integrating high-accuracy PacBio HiFi long-read sequencing and Hi-C chromatin conformation capture technologies and comprehensively annotated with the aid of transcriptomic data. The genome of E. taeniura spans 1.62 Gb, with a scaffold N50 of 206.4 Mb and the longest scaffold reaching 344.4 Mb. The genome exhibits a BUSCO completeness score of 98.2% and contains 22,246 protein-coding genes. Comparative analysis between the assembled genome and three other colubrid species revealed high synteny. In addition, several chromosomal fusion and fission events were observed. This reference genome provides a valuable resource for studying the taxonomy, conservation, and evolutionary history of the widely distributed Elaphe species.
Lianming Du, Jiahao Chen, Qin Liu, Songwen Tan, Peng Guo*. High-quality chromosome-level genome assembly of the snake Pseudoxenodon stejnegeri (Squamata: Colubridae). Scientific Data, 2026, 13: 93.
The taxonomy and evolution of the genus Pseudoxenodon have long been poorly studied, and the paucity of genomic data in Pseudoxenodon critically impedes robust phylogenetic reconstruction and evolutionary analyses. Here, we present a chromosome-level reference genome assembly for P. stejnegeri generated through integrating the PacBio HiFi sequencing, Illumina short-read sequencing and Hi-C scaffolding techniques. The final genome size is 1601.26 Mb, with a scaffold N50 of 203.68 Mb and 97.07% assembled sequences anchored onto 18 pseudo-chromosomes. The BUSCO assessment revealed 97.8% completeness. We predicted 21,678 protein-coding genes, of which 17,531 (80.87%) genes were functionally annotated. Approximately 908.04 Mb repeat sequences were detected, representing 56.71% of the assembled sequences. This high-quality chromosome-level genome provides a valuable genomic resource for future studies on phylogenetics, evolution, and genetics of the genus Pseudoxenodon.
Lianming Du, Dalin Sun, Jiahao Chen, Xinyi Zhou, Kelei Zhao, Qianglin Zeng, Nan Yang*. Pytrf: a python package for finding tandem repeats from genomic sequences. BMC Bioinformatics, 2025, 26: 151.
Background Tandem repeats (TRs) are major sources of genetic variation and important genetic markers. Their expansions are not only involved in gene expression regulation but also associated with many nervous system diseases and cancers. However, there is a lack of an efficient tandem repeat identification tool for seamless integration with larger bioinformatics programs developed with the popular Python language. Results We introduce pytrf, a Python package for identification of both exact and approximate TRs from genomic sequences. It allows seamless embedding into other programs developed by Python or using in Python interactive environment and Jupyter notebooks. It also provides command line tools for assisting users to find tandem repeats from FASTA/Q files. Compared to other tools, the pytrf shows the highest performance in aspect of running time with comparable peak memory usage. Conclusions Pytrf provides simple interfaces and command line tools to facilitate identification of tandem repeats from genomic sequences. Pytrf can easily be installed from PyPI (https://pypi.org/project/pytrf) and the source code is freely available at https://github.com/lmdu/pytrf.
Lianming Du, Jiahao Chen, Dalin Sun, Kelei Zhao, Qianglin Zeng, Nan Yang*. Krait2: a versatile software for microsatellite investigation, visualization and marker development. BMC Genomics, 2025, 26: 72.
Background Microsatellites are highly polymorphic repeat sequences ubiquitously interspersed throughout almost all genomes which are widely used as powerful molecular markers in diverse fields. Microsatellite expansions play pivotal roles in gene expression regulation and are implicated in various neurological diseases and cancers. Although much effort has been devoted to developing efficient tools for microsatellite identification, there is still a lack of a powerful tool for large-scale microsatellite analysis. Results We present Krait2, a user-friendly graphical tool for investigating perfect, imperfect and compound microsatellites from FASTA and FASTQ formatted genomic datasets. Krait2 not only provides features such as primer design, repeat filtering, repeat annotation and statistical analysis, but also offers various output formats to support customized downstream analysis. Moreover, it has capability of analyzing multiple genomes simultaneously and conducting comparative analysis. Conclusions Krait2 is a versatile and easy-to-use software for both novices and experts to identify and analyze microsatellites. The installer and source code are available at https://github.com/lmdu/krait2.
Pei Chen†, Jiangyue Qin†, Helene K Su, Lianming Du*, Qianglin Zeng*. Harmine acts as a quorum sensing inhibitor decreasing the virulence and antibiotic resistance of Pseudomonas aeruginosa. BMC Infectious Diseases, 2024, 24: 760.
Background As antimicrobial resistance (AMR) has become a global health crisis, new strategies against AMR infection are urgently needed. Quorum sensing (QS), responsible for bacterial communication and pathogenicity, is among the targets for anti-virulence drugs that thrive as one of the promising treatments against AMR infection. Methods We identified a natural compound, Harmine, through virtual screening based on three QS receptors of Pseudomonas aeruginosa (P. aeruginosa) and explored the effect of Harmine on QS-controlled and pathogenicity-related phenotypes including pyocyanin production, exocellular protease excretion, biofilm formation, and twitching motility of P. aeruginosa PA14. The protective effect of Harmine on Caenorhabditis elegans (C. elegans) and mice infection models was determined and the synergistic effect of Harmine combined with common antibiotics was explored. The underlaying mechanism of Harmine’s QS inhibitory effect was illustrated by molecular docking analysis, transcriptomic analysis, and target verification assay. Results In vitro results suggested that Harmine possessed QS inhibitory effects on pyocyanin production, exocellular protease excretion, biofilm formation, and twitching motility of P. aeruginosa PA14, and in vivo results displayed Harmine’s protective effect on C. elegans and mice infection models. Intriguingly, Harmine increased susceptibility of both PA14 and clinical isolates of P. aeruginosa to polymyxin B and kanamycin when used in combination. Moreover, Harmine down-regulated a series of QS controlled genes associated with pathogenicity and the underlying mechanism may have involved competitively antagonizing autoinducers’ receptors LasR, RhlR, and PqsR.
Lianming Du, Chaoyue Geng, Qianglin Zeng, Ting Huang, Jie Tang, Yiwen Chu, Kelei Zhao*. Dockey: a modern integrated tool for large-scale molecular docking and virtual screening. Briefings in Bioinformatics, 2023, 24(2): bbad047.
Molecular docking is a structure-based and computer-aided drug design approach that plays a pivotal role in drug discovery and pharmaceutical research. AutoDock is the most widely used molecular docking tool for study of protein–ligand interactions and virtual screening. Although many tools have been developed to streamline and automate the AutoDock docking pipeline, some of them still use outdated graphical user interfaces and have not been updated for a long time. Meanwhile, some of them lack cross-platform compatibility and evaluation metrics for screening lead compound candidates. To overcome these limitations, we have developed Dockey, a flexible and intuitive graphical interface tool with seamless integration of several useful tools, which implements a complete docking pipeline covering molecular sanitization, molecular preparation, paralleled docking execution, interaction detection and conformation visualization. Specifically, Dockey can detect the non-covalent interactions between small molecules and proteins and perform cross-docking between multiple receptors and ligands. It has the capacity to automatically dock thousands of ligands to multiple receptors and analyze the corresponding docking results in parallel. All the generated data will be kept in a project file that can be shared between any systems and computers with the pre-installation of Dockey. We anticipate that these unique characteristics will make it attractive for researchers to conduct large-scale molecular docking without complicated operations, particularly for beginners. Dockey is implemented in Python and freely available at https://github.com/lmdu/dockey.
Lianming Du, Qin Liu, Zhenxin Fan, Jie Tang, Xiuyue Zhang, Megan Price, Bisong Yue*, Kelei Zhao*. Pyfastx: a robust Python package for fast random access to sequences from plain and gzipped FASTA/Q files. Briefings in Bioinformatics, 2021, 22(4): bbaa368.
FASTA and FASTQ are the most widely used biological data formats that have become the de facto standard to exchange sequence data between bioinformatics tools. With the avalanche of next-generation sequencing data, the amount of sequence data being deposited and accessed in FASTA/Q formats is increasing dramatically. However, the existing tools have very low efficiency at random retrieval of subsequences due to the requirement of loading the entire index into memory. In addition, most existing tools have no capability to build index for large FASTA/Q files because of the limited memory. Furthermore, the tools do not provide support to randomly accessing sequences from FASTA/Q files compressed by gzip, which is extensively adopted by most public databases to compress data for saving storage. In this study, we developed pyfastx as a versatile Python package with commonly used command-line tools to overcome the above limitations. Compared to other tools, pyfastx yielded the highest performance in terms of building index and random access to sequences, particularly when dealing with large FASTA/Q files with hundreds of millions of sequences. A key advantage of pyfastx over other tools is that it offers an efficient way to randomly extract subsequences directly from gzip compressed FASTA/Q files without needing to uncompress beforehand. Pyfastx can easily be installed from PyPI (https://pypi.org/project/pyfastx) and the source code is freely available at https://github.com/lmdu/pyfastx.
Lianming Du†, Tao Guo†, Qin Liu, Jing Li, Xiuyue Zhang, Jinchuan Xing, Bisong Yue, Jing Li*, Zhenxin Fan*. MACSNVdb: a high-quality SNV database for interspecies genetic divergence investigation among macaques. Database, 2020, 2020: baaa027.
Macaques are the most widely used non-human primates in biomedical research. The genetic divergence between these animal models is responsible for their phenotypic differences in response to certain diseases. However, the macaque single nucleotide polymorphism resources mainly focused on rhesus macaque (Macaca mulatta), which hinders the broad research and biomedical application of other macaques. In order to overcome these limitations, we constructed a database named MACSNVdb that focuses on the interspecies genetic diversity among macaque genomes. MACSNVdb is a web-enabled database comprising ~74.51 million high-quality non-redundant single nucleotide variants (SNVs) identified among 20 macaque individuals from six species groups (muttla, fascicularis, sinica, arctoides, silenus, sylvanus). In addition to individual SNVs, MACSNVdb also allows users to browse and retrieve groups of user-defined SNVs. In particular, users can retrieve non-synonymous SNVs that may have deleterious effects on protein structure or function within macaque orthologs of human disease and drug-target genes. Besides position, alleles and flanking sequences, MACSNVdb integrated additional genomic information including SNV annotations and gene functional annotations. MACSNVdb will facilitate biomedical researchers to discover molecular mechanisms of diverse responses to diseases as well as primatologist to perform population genetic studies. We will continue updating MACSNVdb with newly available sequencing data and annotation to keep the resource up to date.
Lianming Du, Qin Liu, Kelei Zhao, Jie Tang, Xiuyue Zhang, Bisong Yue*, Zhenxin Fan*. PSMD: An extensive database for pan-species microsatellite investigation and marker development. Molecular Ecology Resources, 2020, 20(1): 283-291.
Microsatellites are widely distributed throughout nearly all genomes which have been extensively exploited as powerful genetic markers for diverse applications due to their high polymorphisms. Their length variations are involved in gene regulation and implicated in numerous genetic diseases even in cancers. Although much effort has been devoted in microsatellite database construction, the existing microsatellite databases still had some drawbacks, such as limited number of species, unfriendly export format, missing marker development, lack of compound microsatellites and absence of gene annotation, which seriously restricted researchers to perform downstream analysis. In order to overcome the above limitations, we developed PSMD (Pan-Species Microsatellite Database, http://big.cdu.edu.cn/psmd/) as a web-based database to facilitate researchers to easily identify microsatellites, exploit reliable molecular markers and compare microsatellite distribution pattern on genome-wide scale. In current release, PSMD comprises 678,106,741 perfect microsatellites and 43,848,943 compound microsatellites from 18,408 organisms, which covered almost all species with available genomic data. In addition to interactive browse interface, PSMD also offers a flexible filter function for users to quickly gain desired microsatellites from large data sets. PSMD allows users to export GFF3 formatted file and CSV formatted statistical file for downstream analysis. We also implemented an online tool for analysing occurrence of microsatellites with user-defined parameters. Furthermore, Primer3 was embedded to help users to design high-quality primers with customizable settings. To our knowledge, PSMD is the most extensive resource which is likely to be adopted by scientists engaged in biological, medical, environmental and agricultural research.
Lianming Du†, Qin Liu†, Fujun Shen, Zhenxin Fan, Rong Hou, Bisong Yue*, Xiuyue Zhang*. Transcriptome analysis reveals immune-related gene expression changes with age in giant panda (Ailuropoda melanoleuca) blood. Aging, 2019, 11(1): 249–262.
The giant panda (Ailuropoda melanoleuca), an endangered species endemic to western China, has long been threatened with extinction that is exacerbated by highly contagious and fatal diseases. Aging is the most well-defined risk factor for diseases and is associated with a decline in immune function leading to increased susceptibility to infection and reduced response to vaccination. Therefore, this study aimed to determine which genes and pathways show differential expression with age in blood tissues. We obtained 210 differentially expressed genes by RNA-seq, including 146 up-regulated and 64 down-regulated genes in old pandas (18-21yrs) compared to young pandas (2-6yrs). We identified ISG15, STAT1, IRF7 and DDX58 as the hub genes in the protein-protein interaction network. All of these genes were up-regulated with age and played important roles in response to pathogen invasion. Functional enrichment analysis indicated that up-regulated genes were mainly involved in innate immune response, while the down-regulated genes were mainly related to B cell activation. These may suggest that the innate immunity is relatively well preserved to compensate for the decline in the adaptive immune function. In conclusion, our findings will provide a foundation for future studies on the molecular mechanisms underlying immune changes associated with ageing.
Lianming Du , Chi Zhang*, Qin Liu , Xiuyue Zhang , Bisong Yue. Krait: an ultrafast tool for genome-wide survey of microsatellites and primer design. Bioinformatics, 2018, 34(4): 681–683.
Summary Microsatellites are found to be related with various diseases and widely used in population genetics as genetic markers. However, it remains a challenge to identify microsatellite from large genome and screen microsatellites for primer design from a huge result dataset. Here, we present Krait, a robust and flexible tool for fast investigation of microsatellites in DNA sequences. Krait is designed to identify all types of perfect or imperfect microsatellites on a whole genomic sequence, and is also applicable to identification of compound microsatellites. Primer3 was seamlessly integrated into Krait so that users can design primer for microsatellite amplification in an efficient way. Additionally, Krait can export microsatellite results in FASTA or GFF3 format for further analysis and generate statistical report as well as plotting. Availability and implementation Krait is freely available at https://github.com/lmdu/krait under GPL2 License, implemented in C and Python, and supported on Windows, Linux and Mac operating systems.