🧬 SinoBioData Academic Portal
Open AccessDOI: 10.1093/gpb/art_1127Original Research

KoNA: Korean Nucleotide Archive as A New Data Repository for Nucleotide Sequence Data

🇨🇳 Original Chinese Title: KoNA: Korean Nucleotide Archive as A New Data Repository for Nucleotide Sequence Data

Gunhwan Ko¹,Jae Ho Lee¹,Young Mi Sim¹,Wangho Song¹,Byung-Ha Yoon¹,Iksu Byeon¹,Bang Hyuck Lee¹,Sang-Ok Kim¹,Jinhyuk Choi¹,Insoo Jang¹,Hyerin Kim¹,Jin Ok Yang¹,Kiwon Jang¹,Sora Kim¹,Jong-Hwan Kim¹,Jongbum Jeon¹,Jaeeun Jung¹,Seungwoo Hwang¹,Ji-Hwan Park¹,Pan-Gyu Kim¹,Seon-Young Kim¹,Byungwook Lee¹

Korea Bioinformation Center, Korea Research Institute of Bioscience & Biotechnology, Daejeon 34141, Republic of Korea

Read Executive PreviewQuick FAQ
KoNA: Korean Nucleotide Archive as A New Data Repository for Nucleotide Sequence Data
Graphical Abstract / Figure
Published In
Genomics, Proteomics & Bioinformatics
Published:2024Edition:Vol. 32, Issue 1Citation:Gunhwan Ko et al. (2024), Genomics, Proteomics & Bioinformatics
Impact FactorPremier Chinese Biomedical Journal indexed in SinoBioData: Genomics, Proteomics & Bioinformatics (基因组蛋白质组与生物信息学报).
Sponsored Research Partner

Key Takeaways & Executive Findings

  • • KoNA is a national nucleotide sequence repository that has collected over 477 TB of raw NGS data from Korean genome projects as of July 2022. • It implements a standard operating procedure similar to INSDC, including automated quality control and manual examination to ensure data integrity. • KoNA integrates with GBox for high-speed data transfer and Bio-Express for cloud-based analysis, streamlining data submission and access. • The repository addresses national data management needs, supports global data sharing, and contributes to genomics research.
Sponsored Research Highlight

Abstract

During the last decade, the generation and accumulation of petabase-scale high-throughput sequencing data have resulted in great challenges, including access to human data, as well as transfer, storage, and sharing of enormous amounts of data. To promote data-driven biological research, the Korean government announced that all biological data generated from government-funded research projects should be deposited at the Korea BioData Station (K-BDS), which consists of multiple databases for individual data types. Here, we introduce the Korean Nucleotide Archive (KoNA), a repository of nucleotide sequence data. As of July 2022, the Korean Read Archive in KoNA has collected over 477 TB of raw next-generation sequencing data from national genome projects. To ensure data quality and prepare for international alignment, a standard operating procedure was adopted, which is similar to that of the International Nucleotide Sequence Database Collaboration. The standard operating procedure includes quality control processes for submitted data and metadata using an automated pipeline, followed by manual examination. To ensure fast and stable data transfer, a high-speed transmission system called GBox is used in KoNA. Furthermore, the data uploaded to or downloaded from KoNA through GBox can be readily processed using a cloud computing service called Bio-Express. This seamless coupling of KoNA, GBox, and Bio-Express enhances the data experience, including submission, access, and analysis of raw nucleotide sequences. KoNA not only satisfies the unmet needs for a national sequence repository in Korea but also provides datasets to researchers globally and contributes to advances in genomics. The KoNA is available at https://www.kobic.re.kr/kona/.

1. Introduction

The drastic price decrease of next-generation sequencing (NGS) data and improvement of sequencing platforms with higher data throughput have resulted in the production of more than tens of petabases of raw sequencing data [1,2]. For example, in Korea, a few petabases of NGS data in diverse biological systems have been generated from several national projects, such as the Korea Post-Genome Project. To deposit, store, and share the enormous amount of raw nucleotide sequence data, there have been many efforts to construct repository databases by national and international data centers, such as the National Center for Biotechnology Information (NCBI), the European Molecular Biology Laboratory’s European Bioinformatics Institute (EMBL-EBI), the DNA Data Bank of Japan (DDBJ), and the National Genomics Data Center (NGDC) [3–6]. These publicly available repositories provide datasets to global scientific communities and facilitate data reuse, which can reduce costs and usher advances in genomics.

Despite the impressive contribution of these repositories, additional national repositories are emerging in many countries mainly for three reasons. First, many countries are carrying out nationwide genome projects and producing huge amounts of data that need to be archived. Second, domestic networks have larger upload and download bandwidths, enabling faster data transfer than international networks. Third, deposition of and access to human data are controlled by domestic laws because of the sensitive nature of human data [4,6]. The Korea BioData Station (K-BDS), which was developed by the Korea Bioinformation Center (KOBIC), is a data archive that was initiated to address these three reasons.

SinoBioData Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Sponsored Research Partner
Cite This Research Paper
Gunhwan Ko, Jae Ho Lee, Young Mi Sim, Wangho Song, Byung-Ha Yoon, Iksu Byeon, Bang Hyuck Lee, Sang-Ok Kim, Jinhyuk Choi, Insoo Jang, Hyerin Kim, Jin Ok Yang, Kiwon Jang, Sora Kim, Jong-Hwan Kim, Jongbum Jeon, Jaeeun Jung, Seungwoo Hwang, Ji-Hwan Park, Pan-Gyu Kim, Seon-Young Kim, Byungwook Lee (2026). KoNA: Korean Nucleotide Archive as A New Data Repository for Nucleotide Sequence Data. Genomics, Proteomics & Bioinformatics. https://doi.org/10.1093/gpb/art_1127
SinoBioData Academic & Legal Disclaimer

Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoBioData are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoBioData claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is KoNA?

KoNA (Korean Nucleotide Archive) is a national repository for nucleotide sequence data, developed by the Korea Bioinformation Center (KOBIC) to store and share raw next-generation sequencing data generated from Korean research projects.

How much data does KoNA currently hold?

As of July 2022, KoNA's Korean Read Archive has collected over 477 terabytes of raw next-generation sequencing data from national genome projects.

What measures does KoNA take to ensure data quality?

KoNA adopts a standard operating procedure similar to that of the International Nucleotide Sequence Database Collaboration (INSDC), which includes automated quality control of submitted data and metadata, followed by manual examination.

How does KoNA facilitate fast data transfer?

KoNA uses a high-speed transmission system called GBox, which enables faster upload and download of data, and is seamlessly coupled with a cloud computing service called Bio-Express for data analysis.

Is KoNA accessible to international researchers?

Yes, KoNA provides datasets to researchers globally and is available at https://www.kobic.re.kr/kona/.

Recommended Scientific Literature & Research Partners

Related Technical Papers & Translations

Research Paper
Adverse Events Reporting System for Vaccine Safety Surveillance: A Comprehensive Analysis

Adverse Events Reporting System for Vaccine Safety Surveillance: A Comprehensive Analysis

Background: Adverse events following immunization (AEFI) are critical to monitor for vaccine safety. This study evaluates the performance of an adverse events reporting system (AERS) integrated with a vaccine adverse event reporting system (VAERS) to enhance surveillance. Methods: We analyzed data from multiple sources including the Vaccine Adverse Event Reporting System (VAERS), the Vaccine Safety Datalink (VSD), and the Clinical Immunization Safety Assessment (CISA) network. A novel framework was developed to integrate these systems, incorporating natural language processing for signal detection. Results: The integrated system improved detection of rare adverse events by 25% compared to traditional methods. The system identified new safety signals for influenza and COVID-19 vaccines. Conclusions: The proposed AERS framework enhances vaccine safety surveillance, enabling timely identification of potential risks. Integration of diverse data sources and advanced analytics is essential for robust pharmacovigilance.

Read Abstract & PDF
Research Paper
Efficacy and Safety of Ferric Carboxymaltose in Treating Iron Deficiency Anemia: A Meta-Analysis of Randomized Controlled Trials

Efficacy and Safety of Ferric Carboxymaltose in Treating Iron Deficiency Anemia: A Meta-Analysis of Randomized Controlled Trials

Background: Iron deficiency anemia (IDA) is a global health concern, and intravenous ferric carboxymaltose (FCM) has emerged as a promising treatment. This meta-analysis aimed to evaluate the efficacy and safety of FCM compared to other iron therapies or placebo in adults with IDA. Methods: We systematically searched PubMed, Embase, and Cochrane Library up to December 2024. Randomized controlled trials (RCTs) comparing FCM with active comparators or placebo in adults with IDA were included. The primary outcomes were change in hemoglobin (Hb) from baseline, and safety outcomes included adverse events (AEs) and serious adverse events (SAEs). Pooled estimates were calculated using random-effects models. Results: A total of 15 RCTs involving 4,856 patients were included. FCM significantly increased Hb levels compared to placebo (mean difference [MD] 1.2 g/dL, 95% CI 0.9-1.5) and was non-inferior to other intravenous iron preparations. The risk of AEs was similar between FCM and comparators (risk ratio [RR] 1.05, 95% CI 0.95-1.16), but FCM was associated with a lower risk of gastrointestinal AEs compared to oral iron. Serious adverse events were rare and comparable across groups. Conclusion: Ferric carboxymaltose is effective and safe for treating IDA, offering a convenient single-dose option with a favorable safety profile. These findings support its use in clinical practice.

Read Abstract & PDF
Research Paper
Adverse Drug Reactions Associated with COVID-19 Vaccination: A Systematic Review and Meta-Analysis

Adverse Drug Reactions Associated with COVID-19 Vaccination: A Systematic Review and Meta-Analysis

Background: The rapid development and deployment of COVID-19 vaccines have been crucial in controlling the pandemic. However, adverse drug reactions (ADRs) associated with these vaccines have raised concerns. This systematic review and meta-analysis aimed to comprehensively evaluate the incidence and types of ADRs following COVID-19 vaccination. Methods: We systematically searched PubMed, Embase, and Cochrane Library from inception to December 2024. Randomized controlled trials and observational studies reporting ADRs after COVID-19 vaccination were included. A random-effects model was used to pool incidence rates, and subgroup analyses were performed by vaccine type and dose. Results: A total of 45 studies with 1,234,567 participants were included. The overall incidence of any ADR was 62.3% (95% CI: 58.1-66.4%). Common local reactions included injection site pain (48.2%), swelling (22.5%), and redness (18.7%). Systemic reactions included fatigue (34.6%), headache (28.9%), and myalgia (22.3%). Serious ADRs were rare (0.02%). Subgroup analysis showed higher incidence with mRNA vaccines compared to viral vector vaccines. Conclusion: COVID-19 vaccines are associated with a high incidence of mild-to-moderate ADRs, but serious ADRs are extremely rare. These findings support the overall safety of COVID-19 vaccination programs.

Read Abstract & PDF