资源与平台

2019新型冠状病毒信息库

展开
  • 1. 国家生物信息中心&中国科学院北京基因组研究所国家基因组科学数据中心, 北京 100101;
    2. 中国科学院北京基因组研究所基因组科学与信息重点实验室,北京 100101
    3. 中国科学院大学,北京 100049

收稿日期: 2020-01-31

  修回日期: 2020-02-07

  网络出版日期: 2020-02-08

基金资助

国家重点研发计划项目(2016YFE0206600);国家重点研发计划项目(2017YFC1201202);中国科学院“十三五”信息化建设专项(XXH13505-05);中国科学院地球大数据先导A类专项(XDA19050302);中国科学院基因组科学数据中心能力建设项目(0202);中国科学院青年创新促进会和中国科学院关键技术人才项目资助

The 2019 novel coronavirus resource

Expand
  • 1. China National Center for Bioinformation & National Genomics Data Center, Beijing Institute of Genomics, Chinese Academy of Sciences, Beijing 100101, China;
    2. CAS Key Laboratory of Genome Sciences and Information, Beijing Institute of Genomics, Chinese Academy of Sciences, Beijing 100101, China
    3. University of Chinese Academy of Sciences, Beijing 100049, China

Received date: 2020-01-31

  Revised date: 2020-02-07

  Online published: 2020-02-08

Supported by

the National Key Research & Development Program of China(2016YFE0206600);the National Key Research & Development Program of China(2017YFC1201202);13th Five-year Informatization Plan of CAS(XXH13505-05);Strategic Priority Research Program of the Chines Academy of Sciences (CAS)(XDA19050302);Capacity building project of genome science data center of Chinese Academy of Sciences(0202);Key Technology Talent Program of the CAS, The Youth Innovation Promotion Association of Chinese Academy of Sciences

摘要

2019年12月在中国武汉开始爆发的新型肺炎已造成全球25个国家/地区的31516人感染、638人死亡(截止2020年2月7日16时),引起该肺炎的病毒被世界卫生组织命名为2019新型冠状病毒(2019-nCoV)。为促进2019-nCoV数据共享应用并及时向全球公众提供病毒的相关信息,国家生物信息中心(CNCB)/国家基因组科学数据中心(NGDC)建立了2019新型冠状病毒信息库(2019nCoVR,https://bigd.big.ac.cn/ncov)。该信息库整合了来自德国全球流感病毒数据库、美国国家生物技术信息中心、深圳(国家)基因库、国家微生物科学数据中心及CNCB/NGDC等机构公开发布的2019-nCoV核苷酸和蛋白质序列数据、元信息、学术文献、新闻动态、科普文章等信息,开展了不同冠状病毒株的基因组序列变异分析并提供可视化展示。同时,2019nCoVR无缝对接CNCB/NGDC的相关数据库,提供新测序病毒株系的基因组原始测序数据、组装后序列的在线汇交、管理与共享、国际数据库同步发布等数据服务。本文对2019nCoVR数据汇交、管理、发布及使用等进行全面阐述,以方便用户了解该信息库各项功能及数据状况,为加速开展病毒的分类溯源、变异演化、快速检测、药物研发以及新型肺炎的精准预防与治疗等研究提供重要基础。

本文引用格式

赵文明, 宋述慧, 陈梅丽, 邹东, 马利娜, 马英克, 李茹姣, 郝丽丽, 李翠萍, 田东梅, 唐碧霞, 王彦青, 朱军伟, 陈焕新, 章张, 薛勇彪, 鲍一明 . 2019新型冠状病毒信息库[J]. 遗传, 2020 , 42(2) : 212 -221 . DOI: 10.16288/j.yczz.20-030

Abstract

An ongoing outbreak of a novel coronavirus infection in Wuhan, China since December 2019 has led to 31,516 infected persons and 638 deaths across 25 countries (till 16:00 on February 7, 2020). The virus causing this pneumonia was then named as the 2019 novel coronavirus (2019-nCoV) by the World Health Organization. To promote the data sharing and make all relevant information of 2019-nCoV publicly available, we construct the 2019 Novel Coronavirus Resource (2019nCoVR, https://bigd.big.ac.cn/ncov). 2019nCoVR features comprehensive integration of genomic and proteomic sequences as well as their metadata information from the Global Initiative on Sharing All Influenza Data, National Center for Biotechnology Information, China National GeneBank, National Microbiology Data Center and China National Center for Bioinformation (CNCB)/National Genomics Data Center (NGDC). It also incorporates a wide range of relevant information including scientific literatures, news, and popular articles for science dissemination, and provides visualization functionalities for genome variation analysis results based on all collected 2019-nCoV strains. Moreover, by linking seamlessly with related databases in CNCB/NGDC, 2019nCoVR offers virus data submission and sharing services for raw sequence reads and assembled sequences. In this report, we provide comprehensive descriptions on data deposition, management, release and utility in 2019nCoVR, laying important foundations in aid of studies on virus classification and origin, genome variation and evolution, fast detection, drug development and pneumonia precision prevention and therapy.

参考文献

[1] WHO. Novel Coronavirus (2019-nCoV). https://www.who.int/emergencies/diseases/novel-coronavirus-2019, 2020.
[2] Xu XT, Chen P, Wang JF, Feng JN, Zhou H, Li X, Zhong W, Hao P . Evolution of the novel coronavirus from the ongoing Wuhan outbreak and modeling of its spike protein for risk of human transmission. Sci China Life Sci, 2020, doi: 10.1007/s11427-020-1637-5.
[3] Zhou P, Yang XL, Wang XG, Hu B, Zhang L, Zhang W, Si HR, Zhu Y, Li B, Huang CL, Chen HD, Chen J, Luo Y, Guo H, Jiang RD, Liu MQ, Chen Y, Shen XR, Wang X, Zheng XS, Zhao K, Chen QJ, Deng F, Liu LL, Yan B, Zhan FX, Wang YY, Xiao GF, Shi ZL . Discovery of a novel coronavirus associated with the recent pneumonia outbreak in humans and its potential bat origin. bioRxiv, 2020, doi: 10.1101/2020.01.22.914952.
[4] Ji W, Wang W, Zhao XF, Zai JJ, Li XG . Homologous recombination within the spike glycoprotein of the newly identified coronavirus may boost cross-species transmission from snake to human. J Med Virol, 2020, doi: 10.1002/jmv.25682.
[5] Dong N, Yang XM, Ye LW, Chen KC, Chan EWC, Yang MS, Chen S . Genomic and protein structure modelling analysis depicts the origin and infectivity of 2019-nCoV, a new coronavirus which caused a pneumonia outbreak in Wuhan, China. bioRxiv, 2020, doi: 10.1101/2020.01.20. 913368.
[6] Benvenuto D, Giovanetti M, Ciccozzi A, Spoto S, Angeletti S, Ciccozzi M . The 2019-new Coronavirus epidemic: evidence for virus evolution. J Med Virol, 2020, doi: 10.1101/2020.01.24.915157.
[7] Chen JY, Shi JS, Qiu DA, Liu C, Li X, Zhao Q, Ruan JS, Gao S . Bioinformatics analysis of the Wuhan 2019 human coronavirus genome. Chin J Bio, 2020, doi: 10.12113/ 202001007.
[7] 陈嘉源, 施劲松, 丘栋安, 刘畅, 李鑫, 赵强, 阮吉寿, 高山 . 武汉2019冠状病毒基因组的生物信息学分析. 生物信息学, 2020, doi: 10.12113/202001007.
[8] Heymann DL . Data sharing and outbreaks: best practice exemplified. Lancet, 2020, doi: 10.1016/S0140-6736(20)30184-7.
[9] Munster VJ , Koopmans M, van Doremalen N, van Riel D, de Wit E. A novel coronavirus emerging in China - key questions for impact assessment. N Engl J Med, 2020, doi: 10.1056/NEJMp2000929.
[10] Sayers EW, Beck J, Brister JR, Bolton EE, Canese K, Comeau DC, Funk K, Ketter A, Kim S, Kimchi A, Kitts PA, Kuznetsov A, Lathrop S, Lu Z , McGarvey K, Madden TL, Murphy TD, O'Leary N, Phan L, Schneider VA, Thibaud-Nissen F, Trawick BW, Pruitt KD, Ostell J. Database resources of the National Center for Biotechnology Information. Nucleic Acids Res, 2020,48:D9-D16.
[11] Shu YL , McCauley J. GISAID: Global initiative on sharing all influenza data-from vision to reality. Euro Surveill, 2017,22(13):30494.
[12] Wang B, Liu F, Zhang EC, Wo CL, Chen J, Qian PY, Lu HR, Zeng WJ, Chen T, Wei JP, Wan Q, Wang R, Xu X . The China National GeneBank—owned by all, completed by all and shared by all. Hereditas(Beijing), 2019,41(8):761-772.
[12] 王博, 刘芳, 张二春, 沃晨亮, 陈振家, 钱璞毅, 卢浩荣, 曾文君, 陈泰, 危金普, 万仟, 王韧, 徐讯 . 国家基因库: 共有、共为、共享. 遗传, 2019,41(8):761-72.
[13] Wu LH, Sun QL, Desmeth P, Sugawara H, Xu ZH , McCluskey K, Smith D, Alexander V, Lima N, Ohkuma M, Robert V, Zhou YG, Li JH, Fan GM, Ingsriswang S, Ozerskaya S, Ma JC. World data centre for microorganisms: an information infrastructure to explore and utilize preserved microbial strains worldwide. Nucleic Acids Res, 2017,45(D1):D611-D618.
[14] Zhang Z, Bao Y, Zhao W, Xiao J, Chen R, Zhang G, Li Y, Zhao G, Pervaiz N, Li R, Gao F, Zhi X, Lu Y, Liu L, He S, Li Q, Yuan C, Ma L, Xiao Y, Wang J, Hao Y, Wang Q, Shang Y, Zhang Y, Yuan N, Song S, Tian F, Sun L, Teng Y, Sun X, Chen H, Xue Y, Zhang Q, Teng X, Huang Z, Wang H, Zhu T, Zhang C, Ma Y, Zhang X, Lin S, Gao Y, Zhou J, Guo J, Liu X, Kang H, Tian D, Gao G, Ling Y, Xu S, Wang P, Zhou H, Niu Y, Ruan C, Lv D, Dong L, Zhu Q, Abbasi AA, Tang Q, Li H, Yao L, Chen M, Gao Q, Cao R, Guo Y, Zhai S, Shi S, Guo AY, Shireen H, Miao YR, Jin JP, Qian Q, Wang Y, Cao J, Duan G, Ning Z, Yu L, Li Z, Du Q, Wu W, Zhou Q, Hu H, Wang G, Wu S, Li CY, Zhao F, Xiong Z, Wang C, Gong Z, Zeng J, Yuan L, Xia X, Sun M, Batool F, Xue H, Sang J, Du Z, Wang X, Lan L, Fang S, Cui Q, Wang Z, Hao L, Liu W, Jiang Z, Zhang H, Raza RZ, Wu Y, Luo H, Zhang YE, Zhu J, Jiang M, Li M, Ying C, Li X, Li C, Zhao Y, Kang Q, Klenk HP, Zheng Y, Yang F, Tang B, Zhang P, Chen X, Zhang L, Zhao L, Tu Y, Chen T, Zou D, Zhang S, Ning W, Niu G, Guo H, Yan J, Shi Y, Sun Y, Pan M, Lu M, Ji P, Peng D, Yuan H . Database resources of the National Genomics Data Center in 2020. Nucleic Acids Res, 2020,48(D1):D24-D33.
[15] Wang YQ, Song FH, Zhu JW, Zhang SS, Yang YD, Chen TT, Tang BX, Dong LL, Ding N, Zhang Q, Bai ZX, Dong XN, Chen HX, Sun MY, Zhai S, Sun YB, Yu L, Lan L, Xiao JF, Fang XD, Lei HX, Zhang Z, Zhao WM . GSA: genome sequence archive. Genomics Proteomics Bioinformatics, 2017,15(1):14-18.
[16] Zhang SS, Chen TT, Zhu JW, Zhou Q, Chen X, Wang YQ, Zhao WM . GSA: genome sequence archive. Hereditas (Beijing), 2018,40(11):1044-1047.
[16] 张思思, 陈婷婷, 朱军伟, 周晴, 陈旭, 王彦青, 赵文明 . GSA: 组学原始数据归档库. 遗传, 2018,40(11):1044-1047.
[17] Zhang YS, Xia L, Sang J, Li M, Liu L, Li MW, Niu GY, Cao JB, Teng XF, Zhou Q, Zhang Z . The BIG Data Center’s database resources. Hereditas(Beijing), 2018,40(11):1039-1043.
[17] 张源笙, 夏琳, 桑健, 李漫, 刘琳, 李萌伟, 牛广艺, 曹佳宝, 滕徐菲, 周晴, 章张 . 生命与健康大数据中心资源. 遗传, 2018,40(11):1039-1043.
[18] Ma YK, Bao YM . Prospects for national biological big data centers. Hereditas(Beijing), 2018,40(11):938-943.
[18] 马英克, 鲍一明 . 国家级生物大数据中心展望. 遗传, 2018,40(11):938-943.
[19] Edgar RC . MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res, 2004,32(5):1792-1797.
[20] Bodenhofer U, Bonatesta E, Horej?-Kainrath C , Hochreiter S. msa: an R package for multiple sequence alignment. Bioinformatics, 2015,31(24):3997-3999.
[21] Donlin MJ . Using the generic genome browser(GBrowse) . Curr Protoc Bioinformatics, 2009, 28(1): 9.9.1-9.9.25.
[22] McLaren W, Gil L, Hunt SE, Riat HS, Ritchie GRS, Thormann A, Flicek P, Cunningham F . The ensembl variant effect predictor. Genome Biol, 2016,17:122.
文章导航

/