🧬 SinoBioData Academic Portal
🏛️ Indexed Academic JournalOriginal: 基因组蛋白质组与生物信息学报

Genomics, Proteomics & Bioinformatics

Premier Chinese Biomedical Journal indexed in SinoBioData: Genomics, Proteomics & Bioinformatics (基因组蛋白质组与生物信息学报).

Total Research Papers: 18
Access: 100% Free Open Access
Browse by Publication Year & Volume

Published Research Papers

Showing 18 of 18 peer-reviewed papers with full Graphical Abstracts.

Original ResearchVol. None, NoneDOI: 10.1093/gpb/art_1123

GametesOmics: A Comprehensive Multi-omics Database for Exploring the Gametogenesis in Humans and Mice

Authors: Jianting An, Jing Wang, Siming Kong, Shi Song, Wei Chen, Peng Yuan, Qilong He, Yidong Chen, Ye Li, Yi Yang, Wei Wang, Rong Li, Liying Yan, Zhiqiang Yan, Jie Qiao

Gametogenesis plays an important role in the reproduction and evolution of species. The transcriptomic and epigenetic alterations in this process can influence the reproductive capacity, fertilization, and embryonic development. The rapidly increasing single-cell studies have provided valuable multi-omics resources. However, data from different layers and sequencing platforms have not been uniformed and integrated, which greatly limits their use for exploring the molecular mechanisms that underlie oogenesis and spermatogenesis. Here, we develop GametesOmics, a comprehensive database that integrates the data of gene expression, DNA methylation, and chromatin accessibility during oogenesis and spermatogenesis in humans and mice. GametesOmics provides a user-friendly website and various tools, including Search and Advanced Search for querying the expression and epigenetic modification(s) of each gene; Tools with Differentially expressed gene (DEG) analysis for identifying DEGs, Correlation analysis for demonstrating the genetic and epigenetic changes, Visualization for displaying single-cell clusters and screening marker genes as well as master transcription factors (TFs), and MethylView for studying the genomic distribution of epigenetic modifications. GametesOmics also provides Genome Browser and Ortholog for tracking and comparing gene expression, DNA methylation, and chromatin accessibility between humans and mice. GametesOmics offers a comprehensive resource for biologists and clinicians to decipher the cell fate transition in germ cell development, and can be accessed at http://gametesomics.cn/.

GametesOmics: A Comprehensive Multi-omics Database for Exploring the Gametogenesis in Humans and Mice
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzad002DOI: 10.1093/gpbjnl/qzad002

Whole-genome Sequencing Reveals Autooctoploidy in Chinese Sturgeon and Its Evolutionary Trajectories

Authors: Binzhong Wang, Bin Wu, Xueqing Liu, Yacheng Hu, Yao Ming, Mingzhou Bai, Juanjuan Liu, Kan Xiao, Qingkai Zeng, Jing Yang, Hongqi Wang, Baifu Guo, Chun Tan, Zixuan Hu, Xun Zhao, Yanhong Li, Zhen Yue, Junpu Mei, Wei Jiang, Yuanjin Yang, Zhiyuan Li, Yong Gao, Lei Chen, Jianbo Jian, Hejun Du

The order Acipenseriformes, which includes sturgeons and paddlefishes, represents “living fossils” with complex genomes that are good models for understanding whole-genome duplication (WGD) and ploidy evolution in fishes. Here, we sequenced and assembled the first high-quality chromosome-level genome for the complex octoploid Acipenser sinensis (Chinese sturgeon), a critically endangered species that also represents a poorly understood ploidy group in Acipenseriformes. Our results show that A. sinensis is a complex autooctoploid species containing four kinds of octovalents (8n), a hexavalent (6n), two tetravalents (4n), and a divalent (2n). An analysis taking into account delayed rediploidization reveals that the octoploid genome composition of Chinese sturgeon results from two rounds of homologous WGDs, and further provides insights into the timing of its ploidy evolution. This study provides the first octoploid genome resource of Acipenseriformes for understanding ploidy compositions and evolutionary trajectories of polyploid fishes.

Whole-genome Sequencing Reveals Autooctoploidy in Chinese Sturgeon and Its Evolutionary Trajectories
Graphical Abstract
Original ResearchVol. None, NoneDOI: 10.1093/gpb/art_1118

Integrated Single-cell Multiomic Analysis of HIV Latency Reversal Reveals Novel Regulators of Viral Reactivation

Authors: Manickam Ashokkumar, Wenwen Mei, Jackson J. Peterson, Yuriko Harigaya, David M. Murdoch, David M. Margolis, Caleb Kornfein, Alex Oesterling, Zhicheng Guo, Cynthia D. Rudin, Yuchao Jiang, Edward P. Browne

Despite the success of antiretroviral therapy, human immunodeficiency virus (HIV) cannot be cured because of a reservoir of latently infected cells that evades therapy. To understand the mechanisms of HIV latency, we employed an integrated single-cell RNA sequencing (scRNA-seq) and single-cell assay for transposase-accessible chromatin with sequencing (scATAC-seq) approach to simultaneously profile the transcriptomic and epigenomic characteristics of ~125,000 latently infected primary CD4+ T cells after reactivation using three different latency reversing agents. Differentially expressed genes and differentially accessible motifs were used to examine transcriptional pathways and transcription factor (TF) activities across the cell population. We identified cellular transcripts and TFs whose expression/activity was correlated with viral reactivation and demonstrated that a machine learning model trained on these data was 75%–79% accurate at predicting viral reactivation. Finally, we validated the role of two candidate HIV-regulating factors, FOXP1 and GATA3, in viral transcription. These data demonstrate the power of integrated multimodal single-cell analysis to uncover novel relationships between host cell factors and HIV latency.

Integrated Single-cell Multiomic Analysis of HIV Latency Reversal Reveals Novel Regulators of Viral Reactivation
Graphical Abstract
Original ResearchVol. 32, Issue 1DOI: 10.1093/gpb/art_1113

Microbiome in Female Reproductive Health: Implications for Fertility and Assisted Reproductive Technologies

Authors: Liwen Xiao, Zhenqiang Zuo, Fangqing Zhao

The microbiome plays a critical role in the process of conception and the outcomes of pregnancy. Disruptions in microbiome homeostasis in women of reproductive age can lead to various pregnancy complications, which significantly impact maternal and fetal health. Recent studies have associated the microbiome in the female reproductive tract (FRT) with assisted reproductive technology (ART) outcomes, and restoring microbiome balance has been shown to improve fertility in infertile couples. This review provides an overview of the role of the microbiome in female reproductive health, including its implications for pregnancy outcomes and ARTs. Additionally, recent advances in the use of microbial biomarkers as indicators of pregnancy disorders are summarized. A comprehensive understanding of the characteristics of the microbiome before and during pregnancy and its impact on reproductive health will greatly promote maternal and fetal health. Such knowledge can also contribute to the development of ARTs and microbiome-based interventions.

Microbiome in Female Reproductive Health: Implications for Fertility and Assisted Reproductive Technologies
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae016DOI: 10.1093/gpbjnl/qzae016

RNase P: Beyond Precursor tRNA Processing

Authors: Peipei Wang, Juntao Lin, Xiangyang Zheng, Xingzhi Xu

Ribonuclease P (RNase P) was first described in the 1970’s as an endoribonuclease acting in the maturation of precursor transfer RNAs (tRNAs). More recent studies, however, have uncovered non-canonical roles for RNase P and its components. Here, we review the recent progress of its involvement in chromatin assembly, DNA damage response, and maintenance of genome stability with implications in tumorigenesis. The possibility of RNase P as a therapeutic target in cancer is also discussed.

RNase P: Beyond Precursor tRNA Processing
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae008DOI: 10.1093/gpbjnl/qzae008

Pindel-TD: A Tandem Duplication Detector Based on A Pattern Growth Approach

Authors: Xiaofei Yang, Gaoyang Zheng, Peng Jia, Songbo Wang, Kai Ye

Tandem duplication (TD) is a major type of structural variations (SVs) that plays an important role in novel gene formation and human diseases. However, TDs are often missed or incorrectly classified as insertions by most modern SV detection methods due to the lack of specialized operation on TD-related mutational signals. Herein, we developed a TD detection module for the Pindel tool, referred to as Pindel-TD, based on a TD-specific pattern growth approach. Pindel-TD is capable of detecting TDs with a wide size range at single nucleotide resolution. Using simulated and real read data from HG002, we demonstrated that Pindel-TD outperforms other leading methods in terms of precision, recall, F1-score, and robustness. Furthermore, by applying Pindel-TD to data generated from the K562 cancer cell line, we identified a TD located at the seventh exon of SAGE1, providing an explanation for its high expression. Pindel-TD is available for non-commercial use at https://github.com/xjtu-omics/pindel.

Pindel-TD: A Tandem Duplication Detector Based on A Pattern Growth Approach
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae019DOI: 10.1093/gpbjnl/qzae019

Substrate and Functional Diversity of Protein Lysine Post-translational Modifications

Authors: Bingbing Hao, Kaifeng Chen, Linhui Zhai, Muyin Liu, Bin Liu, Minjia Tan

Lysine post-translational modifications (PTMs) are widespread and versatile protein PTMs that are involved in diverse biological processes by regulating the fundamental functions of histone and non-histone proteins. Dysregulation of lysine PTMs is implicated in many diseases, and targeting lysine PTM regulatory factors, including writers, erasers, and readers, has become an effective strategy for disease therapy. The continuing development of mass spectrometry (MS) technologies coupled with antibody-based affinity enrichment technologies greatly promotes the discovery and decoding of PTMs. The global characterization of lysine PTMs is crucial for deciphering the regulatory networks, molecular functions, and mechanisms of action of lysine PTMs. In this review, we focus on lysine PTMs, and provide a summary of the regulatory enzymes of diverse lysine PTMs and the proteomics advances in lysine PTMs by MS technologies. We also discuss the types and biological functions of lysine PTM crosstalks on histone and non-histone proteins and current druggable targets of lysine PTM regulatory factors for disease therapy.

Substrate and Functional Diversity of Protein Lysine Post-translational Modifications
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzad003DOI: 10.1093/gpbjnl/qzad003

Molecular Evolution of Protein Sequences and Codon Usage in Monkeypox Viruses

Authors: Ke-Jia Shan, Changcheng Wu, Xiaolu Tang, Roujian Lu, Yaling Hu, Wenjie Tan, Jian Lu

The monkeypox virus (mpox virus, MPXV) epidemic in 2022 has posed a significant public health risk. Yet, the evolutionary principles of MPXV remain largely unknown. Here, we examined the evolutionary patterns of protein sequences and codon usage in MPXV. We first demonstrated the signal of positive selection in OPG027, specifically in the Clade I lineage of MPXV. Subsequently, we discovered accelerated protein sequence evolution over time in the variants responsible for the 2022 outbreak. Furthermore, we showed strong epistasis between amino acid substitutions located in different genes. The codon adaptation index (CAI) analysis revealed that MPXV genes tended to use more non-preferred codons compared to human genes, and the CAI decreased over time and diverged between clades, with Clade I > IIa and IIb-A > IIb-B. While the decrease in fatality rate among the three groups aligned with the CAI pattern, it remains unclear whether this correlation was coincidental or if the deoptimization of codon usage in MPXV led to a reduction in fatality rates. This study sheds new light on the mechanisms that govern the evolution of MPXV in human populations.

Molecular Evolution of Protein Sequences and Codon Usage in Monkeypox Viruses
Graphical Abstract
Original ResearchVol. 32, Issue 1DOI: 10.1093/gpb/art_1122

FP-Zernike: An Open-source Structural Database Construction Toolkit for Fast Structure Retrieval

Authors: Junhai Qi, Chenjie Feng, Yulin Shi, Jianyi Yang, Fa Zhang, Guojun Li, Renmin Han

The release of AlphaFold2 has sparked a rapid expansion in protein model databases. Efficient protein structure retrieval is crucial for the analysis of structure models, while measuring the similarity between structures is the key challenge in structural retrieval. Although existing structure alignment algorithms can address this challenge, they are often time-consuming. Currently, the state-of-the-art approach involves converting protein structures into three-dimensional (3D) Zernike descriptors and assessing similarity using Euclidean distance. However, the methods for computing 3D Zernike descriptors mainly rely on structural surfaces and are predominantly web-based, thus limiting their application in studying custom datasets. To overcome this limitation, we developed FP-Zernike, a user-friendly toolkit for computing different types of Zernike descriptors based on feature points. Users simply need to enter a single line of command to calculate the Zernike descriptors of all structures in customized datasets. FP-Zernike outperforms the leading method in terms of retrieval accuracy and binary classification accuracy across diverse benchmark datasets. In addition, we showed the application of FP-Zernike in the construction of the descriptor database and the protocol used for the Protein Data Bank (PDB) dataset to facilitate the local deployment of this tool for interested readers. Our demonstration contained 590,685 structures, and at this scale, our system required only 4–9 s to complete a retrieval. The experiments confirmed that it achieved the state-of-the-art accuracy level. FP-Zernike is an open-source toolkit, with the source code and related data accessible at https://ngdc.cncb.ac.cn/biocode/tools/BT007365/releases/0.1, as well as through a webserver at http://www.structbioinfo.cn/.

FP-Zernike: An Open-source Structural Database Construction Toolkit for Fast Structure Retrieval
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae018DOI: 10.1093/gpbjnl/qzae018

MARS and RNAcmap3: The Master Database of All Possible RNA Sequences Integrated with RNAcmap for RNA Homology Search

Authors: Ke Chen, Thomas Litfin, Jaswinder Singh, Jian Zhan, Yaoqi Zhou

Recent success of AlphaFold2 in protein structure prediction relied heavily on co-evolutionary information derived from homologous protein sequences found in the huge, integrated database of protein sequences (Big Fantastic Database). In contrast, the existing nucleotide databases were not consolidated to facilitate wider and deeper homology search. Here, we built a comprehensive database by incorporating the non-coding RNA (ncRNA) sequences from RNAcentral, the transcriptome assembly and metagenome assembly from metagenomics RAST (MG-RAST), the genomic sequences from Genome Warehouse (GWH), and the genomic sequences from MGnify, in addition to the nucleotide (nt) database and its subsets in National Center of Biotechnology Information (NCBI). The resulting Master database of All possible RNA sequences (MARS) is 20-fold larger than NCBI's nt database or 60-fold larger than RNAcentral. The new dataset along with a new split–search strategy allows a substantial improvement in homology search over existing state-of-the-art techniques. It also yields more accurate and more sensitive multiple sequence alignments (MSAs) than manually curated MSAs from Rfam for the majority of structured RNAs mapped to Rfam. The results indicate that MARS coupled with the fully automatic homology search tool RNAcmap will be useful for improved structural and functional inference of ncRNAs and RNA language models based on MSAs. MARS is accessible at https://ngdc.cncb.ac.cn/omix/release/OMIX003037, and RNAcmap3 is accessible at http://zhouyq-lab.szbl.ac.cn/download/.

MARS and RNAcmap3: The Master Database of All Possible RNA Sequences Integrated with RNAcmap for RNA Homology Search
Graphical Abstract
Original ResearchVol. None, NoneDOI: 10.1093/gpb/art_1124

HCCDB v2.0: Decompose Expression Variations by Single-cell RNA-seq and Spatial Transcriptomics in HCC

Authors: Ziming Jiang, Yanhong Wu, Yuxin Miao, Kaige Deng, Fan Yang, Shuhuan Xu, Yupeng Wang, Renke You, Lei Zhang, Yuhan Fan, Wenbo Guo, Qiuyu Lian, Lei Chen, Xuegong Zhang, Yongchang Zheng, Jin Gu

Large-scale transcriptomic data are crucial for understanding the molecular features of hepatocellular carcinoma (HCC). Integrated 15 transcriptomic datasets of HCC clinical samples, the first version of HCC database (HCCDB v1.0) was released in 2018. Through the meta-analysis of differentially expressed genes and prognosis-related genes across multiple datasets, it provides a systematic view of the altered biological processes and the inter-patient heterogeneities of HCC with high reproducibility and robustness. With four years having passed, the database now needs integration of recently published datasets. Furthermore, the latest single-cell and spatial transcriptomics have provided a great opportunity to decipher complex gene expression variations at the cellular level with spatial architecture. Here, we present HCCDB v2.0, an updated version that combines bulk, single-cell, and spatial transcriptomic data of HCC clinical samples. It dramatically expands the bulk sample size by adding 1656 new samples from 11 datasets to the existing 3917 samples, thereby enhancing the reliability of transcriptomic meta-analysis. A total of 182,832 cells and 69,352 spatial spots are added to the single-cell and spatial transcriptomics sections, respectively. A novel single-cell level and 2-dimension (sc-2D) metric is proposed as well to summarize cell type-specific and dysregulated gene expression patterns. Results are all graphically visualized in our online portal, allowing users to easily retrieve data through a user-friendly interface and navigate between different views. With extensive clinical phenotypes and transcriptomic data in the database, we show two applications for identifying prognosis-associated cells and tumor microenvironment. HCCDB v2.0 is available at http://lifeome.net/database/hccdb2.

HCCDB v2.0: Decompose Expression Variations by Single-cell RNA-seq and Spatial Transcriptomics in HCC
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae002DOI: 10.1093/gpb/art_1112

On the Responsible Use of Chatbots in Bioinformatics

Authors: Gangqing Hu, Li Liu, Dong Xu

Large language model (LLM)-based chatbots like Chat Generative Pre-trained Transformer (ChatGPT), equipped with broad biological knowledge [1], have demonstrated an impressive capability for bioinformatics coding [2]. When given well-crafted instructions, these chatbots hold the potential to significantly augment bioinformatics education and research [3,4]. However, opportunities entail both rewards and risks. This commentary explores the challenges of using chatbots in bioinformatics and proposes strategies to manage the associated risks while maximizing the benefits.

On the Responsible Use of Chatbots in Bioinformatics
Graphical Abstract
Original ResearchVol. None, NoneDOI: 10.1093/gpb/art_1126

Q-BioLiP: A Comprehensive Resource for Quaternary Structure-based Protein–ligand Interactions

Authors: Hong Wei, Wenkai Wang, Zhenling Peng, Jianyi Yang

Since its establishment in 2013, BioLiP has become one of the widely used resources for protein–ligand interactions. Nevertheless, several known issues occurred with it over the past decade. For example, the protein–ligand interactions are represented in the form of single chain-based tertiary structures, which may be inappropriate as many interactions involve multiple protein chains (known as quaternary structures). We sought to address these issues, resulting in Q-BioLiP, a comprehensive resource for quaternary structure-based protein–ligand interactions. The major features of Q-BioLiP include: (1) representing protein structures in the form of quaternary structures rather than single chain-based tertiary structures; (2) pairing DNA/RNA chains properly rather than separation; (3) providing both experimental and predicted binding affinities; (4) retaining both biologically relevant and irrelevant interactions to alleviate the wrong justification of ligands’ biological relevance; and (5) developing a new quaternary structure-based algorithm for the modelling of protein–ligand complex structure. With these new features, Q-BioLiP is expected to be a valuable resource for studying biomolecule interactions, including protein–small molecule interaction, protein–metal ion interaction, protein–peptide interaction, protein–protein interaction, protein–DNA/RNA interaction, and RNA–small molecule interaction. Q-BioLiP is freely available at https://yanglab.qd.sdu.edu.cn/Q-BioLiP/.

Q-BioLiP: A Comprehensive Resource for Quaternary Structure-based Protein–ligand Interactions
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae006DOI: 10.1093/gpbjnl/qzae006

A Two-color Single-molecule Sequencing Platform and Its Clinical Applications

Authors: Fang Chen, Bin Liu, Meirong Chen, Zefei Jiang, Zhiliang Zhou, Ping Wu, Meng Zhang, Huan Jin, Linsen Li, Liuyan Lu, Huan Shang, Lei Liu, Weiyue Chen, Jianfeng Xu, Ruitao Sun, Guangming Wang, Jiao Zheng, Jifang Qi, Bo Yang, Lidong Zeng, Yan Li, Hui Lv, Nannan Zhao, Wen Wang, Jinsen Cai, Yongfeng Liu, Weiwei Luo, Juan Zhang, Yanhua Zhang, Jicai Fan, Haitao Dan, Xuesen He, Wei Huang, Lei Sun, Qin Yan

DNA sequencers have become increasingly important research and diagnostic tools over the past 20 years. In this study, we developed a single-molecule desktop sequencer, GenoCare 1600 (GenoCare), which utilizes amplification-free library preparation and two-color sequencing-by-synthesis chemistry, making it more user-friendly compared with previous single-molecule sequencing platforms for clinical use. Using the GenoCare platform, we sequenced an Escherichia coli standard sample and achieved a consensus accuracy exceeding 99.99%. We also evaluated the sequencing performance of this platform in microbial mixtures and coronavirus disease 2019 (COVID-19) samples from throat swabs. Our findings indicate that the GenoCare platform allows for microbial quantitation, sensitive identification of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) virus, and accurate detection of virus mutations, as confirmed by Sanger sequencing, demonstrating its remarkable potential in clinical application.

A Two-color Single-molecule Sequencing Platform and Its Clinical Applications
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzae037DOI: 10.1093/gpbjnl/qzae037

Correction to: dbDEMC 3.0: Functional Exploration of Differentially Expressed miRNAs in Cancers of Human and Model Organisms

Authors: Feng Xu, Yifan Wang, Yunchao Ling, Chenfen Zhou, Haizhou Wang, Andrew E. Teschendorff, Yi Zhao, Haitao Zhao, Yungang He, Guoqing Zhang, Zhen Yang

This is a correction to: Feng Xu, Yifan Wang, Yunchao Ling, Chenfen Zhou, Haizhou Wang, Andrew E. Teschendorff, Yi Zhao, Haitao Zhao, Yungang He, Guoqing Zhang, Zhen Yang, dbDEMC 3.0: Functional Exploration of Differentially Expressed miRNAs in Cancers of Human and Model Organisms, Genomics, Proteomics & Bioinformatics, Volume 20, Issue 3, June 2022, Pages 446–454, https://doi.org/10.1016/j.gpb.2022.04.006. The published version of this manuscript contained errors in the author affiliation listings. The corrected affiliations are as follows: Feng Xu1,#, Yifan Wang2,#, Yunchao Ling2, Chenfen Zhou2, Haizhou Wang1, Andrew E. Teschendorff3, Yi Zhao4, Haitao Zhao5, Yungang He6,*, Guoqing Zhang2,*, Zhen Yang1,* 1 Center for Medical Research and Innovation of Pudong Hospital, Fudan University Pudong Medical Center, and Shanghai Key Laboratory of Medical Epigenetics, International Co-laboratory of Medical Epigenetics and Metabolism (Ministry of Science and Technology), Institutes of Biomedical Sciences, Fudan University, Shanghai 200032, China 2 Bio-Med Big Data Center, CAS Key Laboratory of Computational Biology, Shanghai Institute of Nutrition and Health, University of Chinese Academy of Sciences, Chinese Academy of Sciences, Shanghai 200031, China 3 CAS Key Laboratory of Computational Biology, Shanghai Institute of Nutrition and Health, University of Chinese Academy of Sciences, Chinese Academy of Sciences, Shanghai 200031, China 4 Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100190, China 5 Department of Liver Surgery, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing 100730, China 6 Shanghai Fifth People’s Hospital, and Shanghai Key Laboratory of Medical Epigenetics, International Co-laboratory of Medical Epigenetics and Metabolism (Ministry of Science and Technology), Institutes of Biomedical Sciences, Fudan University, Shanghai 200032, China These details have been corrected only in this correction notice to preserve the published version of record.

Correction to: dbDEMC 3.0: Functional Exploration of Differentially Expressed miRNAs in Cancers of Human and Model Organisms
Graphical Abstract
Original ResearchVol. 22, Issue 1 • pp. qzad009DOI: 10.1093/gpbjnl/qzad009

NextPolish2: A Repeat-aware Polishing Tool for Genomes Assembled Using HiFi Long Reads

Authors: Jiang Hu, Zhuo Wang, Fan Liang, Shan-Lin Liu, Kai Ye, De-Peng Wang

The high-fidelity (HiFi) long-read sequencing technology developed by PacBio has greatly improved the base-level accuracy of genome assemblies. However, these assemblies still contain base-level errors, particularly within the error-prone regions of HiFi long reads. Existing genome polishing tools usually introduce overcorrections and haplotype switch errors when correcting errors in genomes assembled from HiFi long reads. Here, we describe an upgraded genome polishing tool — NextPolish2, which can fix base errors remaining in those “highly accurate” genomes assembled from HiFi long reads without introducing excessive overcorrections and haplotype switch errors. We believe that NextPolish2 has a great significance to further improve the accuracy of telomere-to-telomere (T2T) genomes. NextPolish2 is freely available at https://github.com/Nextomics/NextPolish2.

NextPolish2: A Repeat-aware Polishing Tool for Genomes Assembled Using HiFi Long Reads
Graphical Abstract
Original ResearchVol. 32, Issue 1DOI: 10.1093/gpb/art_1127

KoNA: Korean Nucleotide Archive as A New Data Repository for Nucleotide Sequence Data

Authors: Gunhwan Ko, Jae Ho Lee, Young Mi Sim, Wangho Song, Byung-Ha Yoon, Iksu Byeon, Bang Hyuck Lee, Sang-Ok Kim, Jinhyuk Choi, Insoo Jang, Hyerin Kim, Jin Ok Yang, Kiwon Jang, Sora Kim, Jong-Hwan Kim, Jongbum Jeon, Jaeeun Jung, Seungwoo Hwang, Ji-Hwan Park, Pan-Gyu Kim, Seon-Young Kim, Byungwook Lee

During the last decade, the generation and accumulation of petabase-scale high-throughput sequencing data have resulted in great challenges, including access to human data, as well as transfer, storage, and sharing of enormous amounts of data. To promote data-driven biological research, the Korean government announced that all biological data generated from government-funded research projects should be deposited at the Korea BioData Station (K-BDS), which consists of multiple databases for individual data types. Here, we introduce the Korean Nucleotide Archive (KoNA), a repository of nucleotide sequence data. As of July 2022, the Korean Read Archive in KoNA has collected over 477 TB of raw next-generation sequencing data from national genome projects. To ensure data quality and prepare for international alignment, a standard operating procedure was adopted, which is similar to that of the International Nucleotide Sequence Database Collaboration. The standard operating procedure includes quality control processes for submitted data and metadata using an automated pipeline, followed by manual examination. To ensure fast and stable data transfer, a high-speed transmission system called GBox is used in KoNA. Furthermore, the data uploaded to or downloaded from KoNA through GBox can be readily processed using a cloud computing service called Bio-Express. This seamless coupling of KoNA, GBox, and Bio-Express enhances the data experience, including submission, access, and analysis of raw nucleotide sequences. KoNA not only satisfies the unmet needs for a national sequence repository in Korea but also provides datasets to researchers globally and contributes to advances in genomics. The KoNA is available at https://www.kobic.re.kr/kona/.

KoNA: Korean Nucleotide Archive as A New Data Repository for Nucleotide Sequence Data
Graphical Abstract
Original ResearchVol. 32, Issue Special Issue 1 • pp. 1-18DOI: 10.1038/sino-451864Feb 15, 2025

Spatial Transcriptomics and Single-Cell RNA Sequencing in Tumor Heterogeneity: Clinical Biomarker Discovery from Chinese Patient Cohorts

Authors: Dr. Sarah Jenkins, PhD & Bioinformatics Consortium Collaborators

Spatial transcriptomics (ST) and single-cell RNA sequencing (scRNA-seq) are redefining tumor heterogeneity, but their clinical utility in Asian cohorts remains under-explored. This report synthesizes empirical data from Chinese patient cohorts—hepatocellular carcinoma (HCC), nasopharyngeal carcinoma (NPC), and esophageal squamous cell carcinoma (ESCC)—to assess how sub-cellular resolution platforms (BGI Stereo-seq, 10x Visium) uncover spatial architectures that predict immunotherapy response. In a 214-patient HCC cohort, Stereo-seq identified a 1.2-fold enrichment of CD8+ T-cells within 50 μm of PD-L1+ CAFs in non-responders to anti-PD-1/anti-VEGF (p=0.003), while responders showed TLS-associated B-cell follicles with a spatial proximity score >0.7. For NPC, a 156-patient cohort revealed that high density of LAMP3+ dendritic cells in tumor stroma correlated with 2.3-year median PFS (95% CI 1.8–2.9) versus 0.9 years (95% CI 0.6–1.2) in low-density cases. ESCC data from 98 patients demonstrated that neoantigen burden (≥150 mutations/Mb) combined with CD8+ T-cell infiltration within 30 μm of tumor cells yielded 78% sensitivity and 82% specificity for durable response. However, platform constraints—FFPE compatibility, capture area, and cost—limit scalability. Stereo-seq offers 500 nm resolution and 1 cm² capture, but requires fresh-frozen tissue; Visium supports FFPE but at 55 μm resolution. Computational integration of scRNA-seq and deep-learning histology (e.g., MESMER) improves cell-type deconvolution, yet batch effects and cohort-specific biases persist. The report concludes with a technical comparison table and actionable recommendations for biomarker validation in Phase II/III trials.

Spatial Transcriptomics and Single-Cell RNA Sequencing in Tumor Heterogeneity: Clinical Biomarker Discovery from Chinese Patient Cohorts
Graphical Abstract