国家级生物大数据中心展望
收稿日期: 2018-07-02
修回日期: 2018-09-19
网络出版日期: 2018-09-20
基金资助
国家重点研发计划项目(2016YFE0206600);中国科学院“十三五”信息化建设专项(XXH13505-05);中国科学院率先行动“百人计划”项目资助
Prospects for national biological big data centers
Received date: 2018-07-02
Revised date: 2018-09-19
Online published: 2018-09-20
Supported by
Supported by National Key Research and Development Program of China(2016YFE0206600);the 13th Five-year Informatization Plan of Chinese Academy of Sciences(XXH13505-05);the 100-Talent Program of Chinese Academy of Sciences
大数据时代下,科学大数据已经成为科技创新和社会经济发展的新动力。我国是生物数据生产大国,生命大数据是人口健康和国家安全的重要战略资源。面对我国生物数据因存储零散、缺乏系统监管而大量丢失和流失,以及严重依赖国际生物组学大数据中心的局面,亟需从国家层面建设我国自己的生命大数据保存和管理体系。本文以美国NCBI为例介绍了国际生物大数据中心的发展历程及现状,阐明我国建立国家级生物大数据中心的重要性、迫切性、当前历史机遇和发展前景。中国科学院北京基因组研究所生命与健康大数据中心为此做了大量努力,并在数据存储、汇交和转化应用上取得了阶段性成果,以期推进我国生物大数据中心的建设,提高生命科学研究的国际竞争力和影响力。
马英克,鲍一明 . 国家级生物大数据中心展望[J]. 遗传, 2018 , 40(11) : 938 -943 . DOI: 10.16288/j.yczz.18-180
In the era of big data, scientific big data have become the new driving force for both science and technology innovation and social and economic development. China is a powerhouse in generating vast quantities of biological data, which are an essential strategic resource for population health and national security. The current situation of data loss due to the isolated data storage and the lack of systematic data monitoring and management, and the heavy dependency on international biological data centers urgently calls for China’s own life big data storage and management system at the national level. Taking NCBI as an example, this article introduces the development history and present situation of the international biological big data centers. In addition, the importance, urgency, current historical opportunity and prospect of establishing a national biological big data center in China are also expounded in detail. In order to promote the development of the national center and improve China’s international competitiveness and influence in life science research, the BIG Data Center at Beijing Institute of Genomics (BIG), Chinese Academy of Sciences, has taken many efforts on big data deposition, integration and translation and achieved initial progress.
Key words: life and health; big data; national level; big data center
| [1] | International Human Genome Sequencing Consortium. Finishing the euchromatic sequence of the human genome. Nature, 2004,431(7011):931-945. | |||
| [2] | 1000 Genomes Project Consortium, Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA, Hurles ME, McVean GA . A map of human genome variation from population-scale sequencing. Nature, 2010,467(7319):1061-1073. | |||
| [3] | ENCODE Project Consortium. The ENCODE ( ENCyclopedia Of DNA Elements) project. Science, 2004,306(5696):636-640. | |||
| [4] | Parry V . Commit to talks on patient data and public health. Nature, 2017,548(7666):137. | |||
| [5] | Turnbull C, Scott RH, Thomas E, Jones L, Murugaesu N, Pretty FB, Halai D, Baple E, Craig C, Hamblin A, Henderson S, Patch C O'Neill A, Devereau A, Smith K, Martin AR, Sosinsky A, McDonagh EM, Sultana R, Mueller M, Smedley D, Toms A, Dinh L, Fowler T, Bale M, Hubbard T, Rendon A, Hill S, Caulfield MJ, 100 000 Genomes Project. The 100 000 genomes project: bringing whole genome sequencing to the NHS. BMJ, 2018,361:k1687. | |||
| [6] | Cancer Genome Atlas Research Network. Comprehensive genomic characterization defines human glioblastoma genes and core pathways. Nature, 2008,455(7216):1061-1068. | |||
| [7] | Turnbaugh PJ, Ley RE, Hamady M, Fraser-Liggett CM, Knight R, Gordon JI . The human microbiome project. Nature, 2007,449(7164):804-810. | |||
| [8] | NIH HMP Working Group, Peterson J, Garges S, Giovanni M, McInnes P, Wang L, Schloss JA, Bonazzi V, McEwen JE, Wetterstrand KA, Deal C, Baker CC, Di Francesco V, Howcroft TK, Karp RW, Lunsford RD, Wellington CR, Belachew T, Wright M, Giblin C, David H, Mills M, Salomon R, Mullins C, Akolkar B, Begg L, Davis C, Grandison L, Humble M, Khalsa J, Little AR, Peavy H, Pontzer C, Portnoy M, Sayre MH, Starke-Reed P, Zakhari S, Read J, Watson B, Guyer M . The NIH human microbiome project. Genome Res, 2009,19(12):2317-2323. | |||
| [9] | Nagasaki M, Yasuda J, Katsuoka F, Nariai N, Kojima K, Kawai Y, Yamaguchi-Kabata Y, Yokozawa J, Danjoh I, Saito S, Sato Y, Mimori T, Tsuda K, Saito R, Pan X, Nishikawa S, Ito S, Kuroki Y, Tanabe O, Fuse N, Kuriyama S, Kiyomoto H, Hozawa A, Minegishi N, Douglas Engel J, Kinoshita K, Kure S, Yaegashi N ToMMo Japanese Reference Panel Project, Yamamoto M. Rare variant discovery by deep whole-genome sequencing of 1, 070 Japanese individuals. Nat Commun, 2015,6:8018. | |||
| [10] | Gudbjartsson DF, Helgason H, Gudjonsson SA, Zink F, Oddson A, Gylfason A, Besenbacher S, Magnusson G, Halldorsson BV, Hjartarson E, Sigurdsson GT, Stacey SN, Frigge ML, Holm H, Saemundsdottir J, Helgadottir HT, Johannsdottir H, Sigfusson G, Thorgeirsson G, Sverrisson JT, Gretarsdottir S, Walters GB, Rafnar T, Thjodleifsson B, Bjornsson ES, Olafsson S, Thorarinsdottir H, Steingrimsdottir T, Gudmundsdottir TS, Theodors A, Jonasson JG, Sigurdsson A, Bjornsdottir G, Jonsson JJ, Thorarensen O, Ludvigsson P, Gudbjartsson H, Eyjolfsson GI, Sigurdardottir O, Olafsson I, Arnar DO, Magnusson OT, Kong A, Masson G, Thorsteinsdottir U, Helgason A, Sulem P, Stefansson K . Large-scale whole-genome sequencin
/
|