🧬 SinoBioData Academic Portal
🏛️ Indexed Academic JournalOriginal: 基因组蛋白质组与生物信息学报

Genomics, Proteomics & Bioinformatics

Premier Chinese Biomedical Journal indexed in SinoBioData: Genomics, Proteomics & Bioinformatics (基因组蛋白质组与生物信息学报).

Total Research Papers: 18
Access: 100% Free Open Access
Browse by Publication Year & VolumeReset All Filters ✕

Published Research PapersFiltered: Year 2024 • 22 • 1

Showing 10 of 18 peer-reviewed papers with full Graphical Abstracts.

Original ResearchVol. 22, Issue 1 • pp. qzad002DOI: 10.1093/gpbjnl/qzad002

Whole-genome Sequencing Reveals Autooctoploidy in Chinese Sturgeon and Its Evolutionary Trajectories

Authors: Binzhong Wang, Bin Wu, Xueqing Liu, Yacheng Hu, Yao Ming, Mingzhou Bai, Juanjuan Liu, Kan Xiao, Qingkai Zeng, Jing Yang, Hongqi Wang, Baifu Guo, Chun Tan, Zixuan Hu, Xun Zhao, Yanhong Li, Zhen Yue, Junpu Mei, Wei Jiang, Yuanjin Yang, Zhiyuan Li, Yong Gao, Lei Chen, Jianbo Jian, Hejun Du

The order Acipenseriformes, which includes sturgeons and paddlefishes, represents “living fossils” with complex genomes that are good models for understanding whole-genome duplication (WGD) and ploidy evolution in fishes. Here, we sequenced and assembled the first high-quality chromosome-level genome for the complex octoploid Acipenser sinensis (Chinese sturgeon), a critically endangered species that also represents a poorly understood ploidy group in Acipenseriformes. Our results show that A. sinensis is a complex autooctoploid species containing four kinds of octovalents (8n), a hexavalent (6n), two tetravalents (4n), and a divalent (2n). An analysis taking into account delayed rediploidization reveals that the octoploid genome composition of Chinese sturgeon results from two rounds of homologous WGDs, and further provides insights into the timing of its ploidy evolution. This study provides the first octoploid genome resource of Acipenseriformes for understanding ploidy compositions and evolutionary trajectories of polyploid fishes.

Whole-genome Sequencing Reveals Autooctoploidy in Chinese Sturgeon and Its Evolutionary Trajectories
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae016DOI: 10.1093/gpbjnl/qzae016

RNase P: Beyond Precursor tRNA Processing

Authors: Peipei Wang, Juntao Lin, Xiangyang Zheng, Xingzhi Xu

Ribonuclease P (RNase P) was first described in the 1970’s as an endoribonuclease acting in the maturation of precursor transfer RNAs (tRNAs). More recent studies, however, have uncovered non-canonical roles for RNase P and its components. Here, we review the recent progress of its involvement in chromatin assembly, DNA damage response, and maintenance of genome stability with implications in tumorigenesis. The possibility of RNase P as a therapeutic target in cancer is also discussed.

RNase P: Beyond Precursor tRNA Processing
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae008DOI: 10.1093/gpbjnl/qzae008

Pindel-TD: A Tandem Duplication Detector Based on A Pattern Growth Approach

Authors: Xiaofei Yang, Gaoyang Zheng, Peng Jia, Songbo Wang, Kai Ye

Tandem duplication (TD) is a major type of structural variations (SVs) that plays an important role in novel gene formation and human diseases. However, TDs are often missed or incorrectly classified as insertions by most modern SV detection methods due to the lack of specialized operation on TD-related mutational signals. Herein, we developed a TD detection module for the Pindel tool, referred to as Pindel-TD, based on a TD-specific pattern growth approach. Pindel-TD is capable of detecting TDs with a wide size range at single nucleotide resolution. Using simulated and real read data from HG002, we demonstrated that Pindel-TD outperforms other leading methods in terms of precision, recall, F1-score, and robustness. Furthermore, by applying Pindel-TD to data generated from the K562 cancer cell line, we identified a TD located at the seventh exon of SAGE1, providing an explanation for its high expression. Pindel-TD is available for non-commercial use at https://github.com/xjtu-omics/pindel.

Pindel-TD: A Tandem Duplication Detector Based on A Pattern Growth Approach
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae019DOI: 10.1093/gpbjnl/qzae019

Substrate and Functional Diversity of Protein Lysine Post-translational Modifications

Authors: Bingbing Hao, Kaifeng Chen, Linhui Zhai, Muyin Liu, Bin Liu, Minjia Tan

Lysine post-translational modifications (PTMs) are widespread and versatile protein PTMs that are involved in diverse biological processes by regulating the fundamental functions of histone and non-histone proteins. Dysregulation of lysine PTMs is implicated in many diseases, and targeting lysine PTM regulatory factors, including writers, erasers, and readers, has become an effective strategy for disease therapy. The continuing development of mass spectrometry (MS) technologies coupled with antibody-based affinity enrichment technologies greatly promotes the discovery and decoding of PTMs. The global characterization of lysine PTMs is crucial for deciphering the regulatory networks, molecular functions, and mechanisms of action of lysine PTMs. In this review, we focus on lysine PTMs, and provide a summary of the regulatory enzymes of diverse lysine PTMs and the proteomics advances in lysine PTMs by MS technologies. We also discuss the types and biological functions of lysine PTM crosstalks on histone and non-histone proteins and current druggable targets of lysine PTM regulatory factors for disease therapy.

Substrate and Functional Diversity of Protein Lysine Post-translational Modifications
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzad003DOI: 10.1093/gpbjnl/qzad003

Molecular Evolution of Protein Sequences and Codon Usage in Monkeypox Viruses

Authors: Ke-Jia Shan, Changcheng Wu, Xiaolu Tang, Roujian Lu, Yaling Hu, Wenjie Tan, Jian Lu

The monkeypox virus (mpox virus, MPXV) epidemic in 2022 has posed a significant public health risk. Yet, the evolutionary principles of MPXV remain largely unknown. Here, we examined the evolutionary patterns of protein sequences and codon usage in MPXV. We first demonstrated the signal of positive selection in OPG027, specifically in the Clade I lineage of MPXV. Subsequently, we discovered accelerated protein sequence evolution over time in the variants responsible for the 2022 outbreak. Furthermore, we showed strong epistasis between amino acid substitutions located in different genes. The codon adaptation index (CAI) analysis revealed that MPXV genes tended to use more non-preferred codons compared to human genes, and the CAI decreased over time and diverged between clades, with Clade I > IIa and IIb-A > IIb-B. While the decrease in fatality rate among the three groups aligned with the CAI pattern, it remains unclear whether this correlation was coincidental or if the deoptimization of codon usage in MPXV led to a reduction in fatality rates. This study sheds new light on the mechanisms that govern the evolution of MPXV in human populations.

Molecular Evolution of Protein Sequences and Codon Usage in Monkeypox Viruses
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae018DOI: 10.1093/gpbjnl/qzae018

MARS and RNAcmap3: The Master Database of All Possible RNA Sequences Integrated with RNAcmap for RNA Homology Search

Authors: Ke Chen, Thomas Litfin, Jaswinder Singh, Jian Zhan, Yaoqi Zhou

Recent success of AlphaFold2 in protein structure prediction relied heavily on co-evolutionary information derived from homologous protein sequences found in the huge, integrated database of protein sequences (Big Fantastic Database). In contrast, the existing nucleotide databases were not consolidated to facilitate wider and deeper homology search. Here, we built a comprehensive database by incorporating the non-coding RNA (ncRNA) sequences from RNAcentral, the transcriptome assembly and metagenome assembly from metagenomics RAST (MG-RAST), the genomic sequences from Genome Warehouse (GWH), and the genomic sequences from MGnify, in addition to the nucleotide (nt) database and its subsets in National Center of Biotechnology Information (NCBI). The resulting Master database of All possible RNA sequences (MARS) is 20-fold larger than NCBI's nt database or 60-fold larger than RNAcentral. The new dataset along with a new split–search strategy allows a substantial improvement in homology search over existing state-of-the-art techniques. It also yields more accurate and more sensitive multiple sequence alignments (MSAs) than manually curated MSAs from Rfam for the majority of structured RNAs mapped to Rfam. The results indicate that MARS coupled with the fully automatic homology search tool RNAcmap will be useful for improved structural and functional inference of ncRNAs and RNA language models based on MSAs. MARS is accessible at https://ngdc.cncb.ac.cn/omix/release/OMIX003037, and RNAcmap3 is accessible at http://zhouyq-lab.szbl.ac.cn/download/.

MARS and RNAcmap3: The Master Database of All Possible RNA Sequences Integrated with RNAcmap for RNA Homology Search
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae002DOI: 10.1093/gpb/art_1112

On the Responsible Use of Chatbots in Bioinformatics

Authors: Gangqing Hu, Li Liu, Dong Xu

Large language model (LLM)-based chatbots like Chat Generative Pre-trained Transformer (ChatGPT), equipped with broad biological knowledge [1], have demonstrated an impressive capability for bioinformatics coding [2]. When given well-crafted instructions, these chatbots hold the potential to significantly augment bioinformatics education and research [3,4]. However, opportunities entail both rewards and risks. This commentary explores the challenges of using chatbots in bioinformatics and proposes strategies to manage the associated risks while maximizing the benefits.

On the Responsible Use of Chatbots in Bioinformatics
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae006DOI: 10.1093/gpbjnl/qzae006

A Two-color Single-molecule Sequencing Platform and Its Clinical Applications

Authors: Fang Chen, Bin Liu, Meirong Chen, Zefei Jiang, Zhiliang Zhou, Ping Wu, Meng Zhang, Huan Jin, Linsen Li, Liuyan Lu, Huan Shang, Lei Liu, Weiyue Chen, Jianfeng Xu, Ruitao Sun, Guangming Wang, Jiao Zheng, Jifang Qi, Bo Yang, Lidong Zeng, Yan Li, Hui Lv, Nannan Zhao, Wen Wang, Jinsen Cai, Yongfeng Liu, Weiwei Luo, Juan Zhang, Yanhua Zhang, Jicai Fan, Haitao Dan, Xuesen He, Wei Huang, Lei Sun, Qin Yan

DNA sequencers have become increasingly important research and diagnostic tools over the past 20 years. In this study, we developed a single-molecule desktop sequencer, GenoCare 1600 (GenoCare), which utilizes amplification-free library preparation and two-color sequencing-by-synthesis chemistry, making it more user-friendly compared with previous single-molecule sequencing platforms for clinical use. Using the GenoCare platform, we sequenced an Escherichia coli standard sample and achieved a consensus accuracy exceeding 99.99%. We also evaluated the sequencing performance of this platform in microbial mixtures and coronavirus disease 2019 (COVID-19) samples from throat swabs. Our findings indicate that the GenoCare platform allows for microbial quantitation, sensitive identification of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) virus, and accurate detection of virus mutations, as confirmed by Sanger sequencing, demonstrating its remarkable potential in clinical application.

A Two-color Single-molecule Sequencing Platform and Its Clinical Applications
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae037DOI: 10.1093/gpbjnl/qzae037

Correction to: dbDEMC 3.0: Functional Exploration of Differentially Expressed miRNAs in Cancers of Human and Model Organisms

Authors: Feng Xu, Yifan Wang, Yunchao Ling, Chenfen Zhou, Haizhou Wang, Andrew E. Teschendorff, Yi Zhao, Haitao Zhao, Yungang He, Guoqing Zhang, Zhen Yang

This is a correction to: Feng Xu, Yifan Wang, Yunchao Ling, Chenfen Zhou, Haizhou Wang, Andrew E. Teschendorff, Yi Zhao, Haitao Zhao, Yungang He, Guoqing Zhang, Zhen Yang, dbDEMC 3.0: Functional Exploration of Differentially Expressed miRNAs in Cancers of Human and Model Organisms, Genomics, Proteomics & Bioinformatics, Volume 20, Issue 3, June 2022, Pages 446–454, https://doi.org/10.1016/j.gpb.2022.04.006. The published version of this manuscript contained errors in the author affiliation listings. The corrected affiliations are as follows: Feng Xu1,#, Yifan Wang2,#, Yunchao Ling2, Chenfen Zhou2, Haizhou Wang1, Andrew E. Teschendorff3, Yi Zhao4, Haitao Zhao5, Yungang He6,*, Guoqing Zhang2,*, Zhen Yang1,* 1 Center for Medical Research and Innovation of Pudong Hospital, Fudan University Pudong Medical Center, and Shanghai Key Laboratory of Medical Epigenetics, International Co-laboratory of Medical Epigenetics and Metabolism (Ministry of Science and Technology), Institutes of Biomedical Sciences, Fudan University, Shanghai 200032, China 2 Bio-Med Big Data Center, CAS Key Laboratory of Computational Biology, Shanghai Institute of Nutrition and Health, University of Chinese Academy of Sciences, Chinese Academy of Sciences, Shanghai 200031, China 3 CAS Key Laboratory of Computational Biology, Shanghai Institute of Nutrition and Health, University of Chinese Academy of Sciences, Chinese Academy of Sciences, Shanghai 200031, China 4 Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100190, China 5 Department of Liver Surgery, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing 100730, China 6 Shanghai Fifth People’s Hospital, and Shanghai Key Laboratory of Medical Epigenetics, International Co-laboratory of Medical Epigenetics and Metabolism (Ministry of Science and Technology), Institutes of Biomedical Sciences, Fudan University, Shanghai 200032, China These details have been corrected only in this correction notice to preserve the published version of record.

Correction to: dbDEMC 3.0: Functional Exploration of Differentially Expressed miRNAs in Cancers of Human and Model Organisms
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzad009DOI: 10.1093/gpbjnl/qzad009

NextPolish2: A Repeat-aware Polishing Tool for Genomes Assembled Using HiFi Long Reads

Authors: Jiang Hu, Zhuo Wang, Fan Liang, Shan-Lin Liu, Kai Ye, De-Peng Wang

The high-fidelity (HiFi) long-read sequencing technology developed by PacBio has greatly improved the base-level accuracy of genome assemblies. However, these assemblies still contain base-level errors, particularly within the error-prone regions of HiFi long reads. Existing genome polishing tools usually introduce overcorrections and haplotype switch errors when correcting errors in genomes assembled from HiFi long reads. Here, we describe an upgraded genome polishing tool — NextPolish2, which can fix base errors remaining in those “highly accurate” genomes assembled from HiFi long reads without introducing excessive overcorrections and haplotype switch errors. We believe that NextPolish2 has a great significance to further improve the accuracy of telomere-to-telomere (T2T) genomes. NextPolish2 is freely available at https://github.com/Nextomics/NextPolish2.

NextPolish2: A Repeat-aware Polishing Tool for Genomes Assembled Using HiFi Long Reads
Graphical Abstract