[an error occurred while processing this directive]
Resource and Platform

CNGBdb: China National GeneBank DataBase

Expand
  • 1. China National GeneBank, Shenzhen 518120, China
    2. BGI-Shenzhen, Shenzhen 518083, China
    3. MGI-Shenzhen, Shenzhen 518083, China
    4. Guangdong Provincial Key Laboratory of Genome Read and Write, Shenzhen 518120, China

Received date: 2020-03-23

  Revised date: 2020-05-23

  Online published: 2020-06-01

Supported by

Supported by Guangdong Provincial Key Laboratory of Genome Read and Write No(2017B030301011)

Abstract

China National GeneBank DataBase (CNGBdb) is a data platform aiming to systematically archiving and sharing of multi-omics data in life science. As the service portal of Bio-informatics Data Center of the core structure, namely, "Three Banks and Two Platforms" of China National GeneBank (CNGB), CNGBdb has the advantages of rich sample resources, data resources, cooperation projects, powerful data computation and analysis capabilities. With the advent of high throughput sequencing technologies, research in life science has entered the big data era, which is in the need of closer international cooperation and data sharing. With the development of China's economy and the increase of investment in life science research, we need to establish a national public platform for data archiving and sharing in life science to promote the systematic management, application and industrial utilization. Currently, CNGBdb can provide genomic data archiving, information search engines, data management and data analysis services. The data schema of CNGBdb has covered projects, samples, experiments, runs, assemblies, variations and sequences. Until May 22, 2020, CNGBdb has archived 2176 research projects and more than 2221 TB sequencing data submitted by researchers globally. In the future, CNGBdb will continue to be dedicated to promoting data sharing in life science research and improving the service capability. CNGBdb website is: https://db.cngb.org/.

Cite this article

Fengzhen Chen, Lijin You, Fan Yang, Lina Wang, Xueqin Guo, Fei Gao, Cong Hua, Cong Tan, Lin Fang, Riqiang Shan, Wenjun Zeng, Bo Wang, Ren Wang, Xun Xu, Xiaofeng Wei . CNGBdb: China National GeneBank DataBase[J]. Hereditas(Beijing), 2020 , 42(8) : 799 -809 . DOI: 10.16288/j.yczz.20-080

References

[1] Wang B, Liu F, Zhang EC, Wo CL, Chen J, Qian PY, Lu HR, Zeng WJ, Chen T, Wei JP, Wan Q, Wang R, Xu X . The China National GeneBank─owned by all, completed by all and shared by all. Hereditas(Beijing), 2019,41(8):761-772.
[1] 王博, 刘芳, 张二春, 沃晨亮, 陈振家, 钱璞毅, 卢浩荣, 曾文君, 陈泰, 危金普, 万仟, 王韧, 徐讯 . 国家基因库: 共有、共为、共享. 遗传, 2019,41(8):761-772.
[2] Clarke L, Fairley S, Zheng-Bradley X, Streeter I, Perry E, Lowy E, Tassé AM, Flicek P . The international Genome sample resource (IGSR): A worldwide collection of genome variation incorporating the 1000 Genomes Project data. Nucleic Acids Res, 2017,45(D1):D854-D859.
[3] Consortium ICG . International network of cancer genome projects. Nature, 2010,464(7291):993-938.
[4] Yu J, Hu SN, Wang J, Wong GKS, Li SG, Liu B, Deng YJ, Dai L, Zhou Y, Zhang XQ, Cao ML, Liu J, Sun JD, Tang JB, Chen YJ, Huang XB, Lin W, Ye C, Tong W, Cong LJ, Geng JN, Han YJ, Li L, Li W, Hu GQ, Huang XG, Li WJ, Li J, Liu ZW, Li L, Liu JP, Qi QH, Liu JS, Li L, Li T, Wang XJ, Lu H, Wu TT, Zhu M, Ni PX, Han H, Dong W, Ren XY, Feng XL, Cui P, Li XR, Wang H, Xu X, Zhai WX, Xu Z, Zhang JS, He SJ, Zhang JG, Xu JC, Zhang KL, Zheng XW, Dong JH, Zeng WY, Tao L, Ye J, Tan J, Ren XD, Chen XW, He J, Liu DF, Tian W, Tian CG, Xia HG, Bao QY, Li G, Gao H, Cao T, Wang J, Zhao WM, Li P, Chen W, Wang XD, Zhang Y, Hu JF, Wang J, Liu S, Yang G, Zhang GY, Xiong YQ, Li ZJ, Mao L, Zhou CS, Zhu Z, Chen RS, Hao BL, Zheng WM, Chen SY, Guo W, Li GJ, Liu SQ, Tao M, Wang J, Zhu LH, Yuan LP, Yang HM . A draft sequence of the rice genome (Oryza sativa L. ssp. indica). Science, 2002,296(5565):79-92.
[5] International RGSP . The map-based sequence of the rice genome. Nature, 2005,436(7052):793-800.
[6] RGP. The 3,000 rice genomes project. GigaScience, 2014,3:7.
[7] Milner SG, Jost M, Taketa S, Mazón ER, Himmelbach A, Oppermann M, Weise S, Knüpffer H, Basterrechea M, K?nig P, Schüler D, Sharma R, Pasam RK, Rutten T, Guo GG, Xu DD, Zhang J, Herren G, Müller T, Krattinger SG, Keller B, Jiang Y, González MY, Zhao YS, Habeku? A, F?rber S, Ordon F, Lange M, B?rner A, Graner A, Reif JC, Scholz U, Mascher M, Stein N . Genebank genomics highlights the diversity of a global barley collection. Nat Genet, 2019,51(2):319-326.
[8] Sayers EW, Cavanaugh M, Clark K, Ostell J, Pruitt KD, Karsch-Mizrachi I . GenBank. Nucleic Acids Res, 2019,47(D1):D94-D99.
[9] Madeira F, Park YM, Lee J, Buso N, Gur T, Madhusoodanan N, Basutkar P, Tivey ARN, Potter SC, Finn RD, Lopez R . The EMBL-EBI search and sequence analysis tools APIs in 2019. Nucleic Acids Res, 2019,47(W1):W636-W641.
[10] Kodama Y, Mashima J, Kosuge T, Ogasawara O . DDBJ update: the Genomic Expression Archive (GEA) for functional genomics data. Nucleic Acids Res, 2019,47(D1):D69-D73.
[11] Rigden DJ, Fernández XM . The 2018 Nucleic Acids Research database issue and the online molecular biology database collection. Nucleic Acids Res, 2018,46(D1):D1-D7.
[12] Members SIB . The SIB Swiss Institute of Bioinformatics’ resources: focus on curated databases. Nucleic Acids Res, 2016,44(D1):D27-D37.
[13] Kanehisa M, Furumichi M, Tanabe M, Sato Y, Morishima K . KEGG: new perspectives on genomes, pathways, diseases and drugs. Nucleic Acids Res, 2017,45(D1):D353-D361.
[14] Cochrane G, Karsch-Mizrachi I, Takagi T , International Nucleotide Sequence Database Collaboration. The international nucleotide sequence database collaboration. Nucleic Acids Res, 2016,46(D1):D48-D51.
[15] Wang J, Wang W, Li RQ, Li YR, Tian G, Goodman L, Fan W, Zhang JQ, Li J, Zhang JB, Guo TR, Feng BX, Li H, Lu Y, Fang XD, Liang HQ, Du ZL, Li D, Zhao YQ, Hu YJ, Yang ZZ, Zheng HC, Hellmann I, Inouye M, Pool J, Yi X, Zhao J, Duan JJ, Zhou Y, Qin JJ, Ma LJ, Li GQ, Yang ZT, Zhang GJ, Yang B, Yu C, Liang F, Li WJ, Li SC, Li DW, Ni PX, Ruan J, Li QB, Zhu HM, Liu DY, Lu ZK, Li N, Guo GW, Zhang JG, Ye J, Fang L, Hao Q, Chen Q, Liang Y, Su YY, San A, Ping C, Yang S, Chen F, Li L, Zhou K, Zheng HK, Ren YY, Yang L, Gao Y, Yang GH, Li Z, Feng XL, Kristiansen K, Wong GKS, Nielsen R, Durbin R, Bolund L, Zhang XQ, Li SG, Yang HM, Wang J . The diploid genome sequence of an Asian individual. Nature, 2008,456(7218):60-65.
[16] Li RQ, Fan W, Tian G, Zhu HM, He L, Cai J, Huang QF, Cai QL, Li B, Bai YQ, Zhang ZH, Zhang YP, Wang W, Li J, Wei FW, Li H, Jian M, Li JW, Zhang ZL, Nielsen R, Li DW, Gu WJ, Yang ZT, Xuan ZL, Ryder OA, Leung FCC, Zhou Y, Cao JJ, Sun X, Fu YG, Fang XD, Guo XS, Wang B, Hou R, Shen FJ, Mu B, Ni PX, Lin RM, Qian WB, Wang GD, Yu C, Nie WH, Wang JH, Wu ZG, Liang HQ, Min JM, Wu Q, Cheng SF, Ruan J, Wang MW, Shi ZB, Wen M, Liu BH, Ren XL, Zheng HS, Dong D, Cook K, Shan G, Zhang H, Kosiol C, Xie XY, Lu ZH, Zheng HC, Li YR, Steiner CC, Tsan-Yuk Lam T, Lin SY, Zhang QH, Li GQ, Tian J, Gong TM, Liu HD, Zhang DJ, Fang L, Ye C, Zhang JB, Hu WB, Xu AL, Ren YY, Zhang GJ, Bruford MW, Li QB, Ma LJ, Guo YR, An N, Hu YJ, Zheng Y, Shi YY, Li ZQ, Liu Q, Chen YL, Zhao J, Qu N, Zhao SC, Tian F, Wang XL, Wang HY, Xu LZ, Liu X, Vinar T, Wang YJ, Lam TW, Yiu SM, Liu SP, Zhang HM, Li DS, Huang Y, Wang X, Yang GH, Jiang Z, Wang JY, Qin N, Li L, Li JX, Bolund L, Kristiansen K, Wong GKS, Olson M, Zhang XQ, Li SG, Yang HM, Wang J, Wang J. The sequence and de novo assembly of the giant panda genome. Nature, 2010,463(7279):311-317.
[17] Members NGDC . Database resources of the national genomics data center in 2020. Nucleic Acids Res, 2020,48(D1):D24-D33.
[18] Ma YK, Bao YM . Prospects for national biological big data centers. Hereditas(Beijing), 2018,40(11):938-943.
[18] 马英克, 鲍一明 . 国家级生物大数据中心展望. 遗传, 2018,40(11):938-943.
[19] Wang YQ, Song FH, Zhu JW, Zhang SS, Yang YD, Chen TT, Tang BX, Dong LL, Ding N, Zhang Q, Bai ZX, Dong XN, Chen HX, Sun MY, Zhai S, Sun YB, Yu L, Lan L, Xiao JF, Fang XD, Lei HX, Zhang Z, Zhao WM . GSA: genome sequence archive. Genomics Proteomics Bioinformatics, 2017,15(1):14-18.
[20] Zhang YS, Xia L, Sang J, Li M, Liu L, Li MW, Niu GY, Cao JB, Teng XF, Zhou Q, Zhang, Z. The BIG Data Center's database resources. Hereditas(Beijing), 2018,40(11):1039-1043.
[20] 张源笙, 夏琳, 桑健, 李漫, 刘琳, 李萌伟, 牛广艺, 曹佳宝, 滕徐菲, 周晴, 章张 . 生命与健康大数据中心资源. 遗传, 2018,40(11):1039-1043.
[21] Zhang SS, Chen TT, Zhu JW, Zhou Q, Chen X, Wang YQ, Zhao WM . GSA: genome sequence archive. Hereditas (Beijing), 2018,40(11):1044-1047.
[21] 张思思, 陈婷婷, 朱军伟, 周晴, 陈旭, 王彦青, 赵文明 . GSA: 组学原始数据归档库. 遗传, 2018,40(11):1044-1047.
[22] Shi WY, Qi HY, Sun QL, Fan GM, Liu SJ, Wang J, Zhu BL, Liu HW, Zhao FQ, Wang XC, Hu XX, Li W, Liu J, Tian Y, Wu LH, Ma JC,. gcMeta: a Global Catalogue of Metagenomics platform to support the archiving, standardization and analysis of microbiome data. Nucleic Acids Res, 2019,47(D1):D637-D648.
[23] Wu LH, Sun QL, Sugawara H, Yang S, Zhou YG, McCluskey K, Vasilenko A, Suzuki KI, Ohkuma M, Lee Y, Robert V, Ingsriswang S, Guissart F, Philippe D, Ma JC. Global catalogue of microorganisms (gcm): a comprehensive database and information retrieval, analysis, and visualization system for microbial resources. BMC Genomics, 2013,14:933.
[24] Zhang GJ . Bird sequencing project takes off. Nature, 2015,522(7554):34.
[25] Fan GY, Song Y, Huang XY, Yang LD, Zhang SY, Zhang MQ, Yang XW, Chang Y, Zhang H, Li YX, Liu SS, Yu LL, Seim I, Feng CG, Wang W, Wang K, Wang J, Xu X, Yang HM, Chen NS, Liu X, He SP . Initial data release and announcement of the Fish10K: Fish 10,000 Genomes Project. bioRxiv, 2019,787028.
[26] Initiative OTPT . One thousand plant transcriptomes and the phylogenomics of green plants. Nature, 2019,574:679-685.
[27] Paskin N . Digital object identifier (DOI?) system. Encyclopedia of Library and Information Sciences, 2010,3:1586-1592.
[28] Smigielski EM, Sirotkin K, Ward M, Sherry ST,. dbSNP: a database of single nucleotide polymorphisms. Nucleic Acids Res, 2000,28(1):352-355.
[29] Landrum MJ, Lee JM, Benson M, Brown G, Chao C, Chitipiralla S, Gu BS, Hart J, Hoffman D, Hoover J, Jang WH, Katz KK, Ovetsky M, Riley G, Sethi A, Tully R, Villamarin-Salomon R, Rubinstein W, Maglott DR . ClinVar: public archive of interpretations of clinically relevant variants. Nucleic Acids Res, 2016,44(D1):D862-D868.
[30] Consortium U . UniProt: a worldwide hub of protein knowledge. Nucleic Acids Res, 2019,47(D1):D506-D515.
[31] Pruitt KD, Tatusova T, Maglott DR . NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins. Nucleic Acids Res, 2007,35:D61-D65.
[32] Barrett T, Clark K, Gevorgyan R, Gorelenkov V, Gribov E, Karsch-Mizrachi I, Kimelman M, Pruitt KD, Resenchuk S, Tatusova T, Yaschenko E, Ostell J . BioProject and BioSample databases at NCBI: facilitating capture and organization of metadata. Nucleic Acids Res, 2012,40(D1):D57-D63.
[33] Kodama Y, Shumway M, Leinonen R . The Sequence Read Archive: explosive growth of sequencing data. Nucleic Acids Res, 2012,40:D54-D56.
[34] Kitts PA, Church DM, Thibaud-Nissen F, Choi J, Hem V, Sapojnikov V, Smith RG, Tatusova T, Xiang C, Zherikov A, DiCuccio M, Murphy TD, Pruitt KD, Kimchi A. Assembly: a resource for assembled genomes at NCBI. Nucleic Acids Res, 2016,44:D73-D80.
[35] Gormley C, Tong Z . Elasticsearch: the definitive guide: a distributed real-time search and analytics engine. “O'Reilly Media, Inc.”, 2015.
[36] Federhen S . The NCBI taxonomy database. Nucleic Acids Res, 2012,40:D136-D143.
[37] Marc DT, Khairat SS,. Medical Subject Headings(MeSH) for indexing and retrieving open-source healthcare data. Stud Health Technol Inform, 2014,202:157-160.
Outlines

/