GSA:组学原始数据归档库
收稿日期: 2018-06-29
修回日期: 2018-09-06
网络出版日期: 2018-10-09
基金资助
国家重点研发计划“国家生物信息平台支撑技术项目”和“精准医学项目”(2017YFC1201200);国家重点研发计划“国家生物信息平台支撑技术项目”和“精准医学项目”(2016YFC0901603);中国科学院战略性先导科技专项基金项目(XDB13040500);中国科学院战略性先导科技专项基金项目(XDA08020102);国家自然科学基金项目(91731304);中国科学院关键技术人才基金项目和中国科学院“十三五”信息化建设专项(XXH13505-05)
GSA: Genome Sequence Archive
Received date: 2018-06-29
Revised date: 2018-09-06
Online published: 2018-10-09
Supported by
[Supported by the National Key R&D Program of China(2017YFC1201200);[Supported by the National Key R&D Program of China(2016YFC0901603);the Strategic Priority Research Program of the Chinese Academy of Sciences(XDB13040500);the Strategic Priority Research Program of the Chinese Academy of Sciences(XDA08020102);the National Natural Science Foundation of China(91731304);Key Technology Talent Program of the Chinese Academy of Sciences and the 13th Five-year Informatization Plan of Chinese Academy of Sciences(XXH13505-05)
生命科学的发展已进入组学大数据时代,然而我国至今尚未形成公共数据库存储体系。为弥补国内空白,组学原始数据归档库(Genome Sequence Archive, GSA, http://bigd.big.ac.cn/gsa)系统遵循国际核苷酸序列数据联盟(International Nucleotide Sequence Database Collaboration,INSDC)相关数据库建设标准,广泛收集各类生命组学原始数据。自2015年底上线运行以来,已获得了包括Cell、Nature、PNAS、GPB等30余个国内外期刊的认可,收录的数据量呈显著增长趋势,提供的数据服务受到国内外广大科研人员的认可。GSA有效缓解了当前我国生命组学数据汇交、存储与共享困难的问题,为我国国家生物信息中心的建设奠定了坚实基础。本文对目前GSA数据汇交、审核、发布与管理等机制进行了深入阐述,以方便用户了解GSA的各项功能,提供更高效的数据服务。
关键词: 组学原始数据归档库(GSA); 组学大数据; 数据汇交; 数据共享
张思思,陈婷婷,朱军伟,周晴,陈旭,王彦青,赵文明 . GSA:组学原始数据归档库[J]. 遗传, 2018 , 40(11) : 1044 -1047 . DOI: 10.16288/j.yczz.18-178
The Genome Sequence Archive (GSA), a new data repository for raw sequence reads in China, has been developed in compliance with the International Nucleotide Sequence Database Collaboration (INSDC) standards. It supports data generated from a variety of sequencing platforms ranging from Sanger sequencing to single-cell sequencing and provides data storing and sharing services freely for worldwide scientific communities. Since it went online in late 2015, GSA has archived more than 500 TB data and been acknowledged by many high-profile journals, including Cell, Nature, PNAS, GPB, etc. Focusing on omics data submission, storing and sharing typically for Chinese users, GSA promotes the initiative of the National Bioinformatics Center of China. This paper introduces the specifies of GSA as data collection, curation, management and exchange to facilitate users to understand and use GSA database.
Key words: Genome Sequence Archive (GSA); omics data; data submission; data sharing
| [1] | NCBI Resource Coordinators . Database resources of the National Center for Biotechnology Information. Nucleic Acids Res, 2018,46(Database issue):D8-D13. |
| [2] | Silvester N, Alako B, Amid C, Cerdeño-Tarrága A, Clarke L, Cleland I, Harrison PW, Jayathilaka S, Kay S, Keane T, Leinonen R, Liu X, Martínez-Villacorta J, Menchi M, Reddy K, Pakseresht N, Rajan J, Rossello M, Smirnov D, Toribio AL, Vaughan D, Zalunin V, Cochrane G . The European Nucleotide Archive in 2017. Nucleic Acids Res, 2018,46(Database issue):D36-D40. |
| [3] | Mashima J, Kodama Y, Kosuge T, Fujisawa T, Katayama T, Nagasaki H, Okuda Y, Kaminuma E, Ogasawara O, Okubo K, Nakamura Y, Takagi T . DNA Data Bank of Japan (DDBJ) progress report. Nucleic Acids Res, 2016,44(Database issue):D51-57. |
| [4] | Karsch-Mizrachi I, Takagi T, Cochrane G , Yang YG.cleotide Sequence Database Collaboration. The International Nucleotide Sequence Database Collaboration. Nucleic Acids Research, 2018,46((Database issue):D48-D51. |
| [5] | BIG Data Center Members . The BIG Data Center: from deposition to integration to translation. Nucleic Acids Res, 2017,45(Database issue):D18-D24. |
| [6] | BIG Data Center Members . Database resources of the BIG Data Center in 2018. Nucleic Acids Res, 2018,46(Database issue):D14-D20. |
| [7] | Wang Y, Song F, Zhu J, Zhang S, Yang Y, Chen T, Tang B, Dong L, Ding N, Zhang Q, Bai Z, Dong X, Chen H, Sun M, Zhai S, Sun Y, Yu L, Lan L, Xiao J, Fang X, Lei H, Zhang Z, Zhao W . GSA: Genome Sequence Archive. Genom Proteom Bioinform, 2017,15(1):14-18. |
/
| 〈 |
|
〉 |