研究报告

基于序列相似性和Z曲线方法重注释原核生物蛋白编码基因

展开
  • 1. 电子科技大学生命科学与技术学院,成都 611731
    2. 西南医科大学基础医学院,分子生物与生物化学教研室,泸州 646000
刘硕,在读博士研究生,专业方向:微生物基因组学。E-mail: liushuo20022020@gmail.com

收稿日期: 2020-02-20

  修回日期: 2020-05-11

  网络出版日期: 2020-05-28

基金资助

电子科技大学理科实力提升计划项目资助编号(Y0301902610100202)

Comprehensive re-annotation of protein-coding genes for prokaryotic genomes by Z-curve and similarity-based methods

Expand
  • 1. School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu 611731, China
    2. Department of Biochemistry and Molecular Biology, School of Basic Medicine, Southwest Medical University, Luzhou 646000,China

Received date: 2020-02-20

  Revised date: 2020-05-11

  Online published: 2020-05-28

Supported by

Supported by Science Strength Improvement Plan of University of Electronic Science and Technology of China No(Y0301902610100202)

摘要

随着测序技术的不断发展,产生了海量的基因组测序数据,极大地丰富了公共遗传数据资源。同时为了应对大量基因组数据的产生,基因组比较和注释算法、工具不断更新,使得联合多种注释工具得到更准确的蛋白编码基因的注释信息成为可能。目前公共数据库的原核生物基因组测序和装配有些是10多年前的,存在大量预测的功能未知的编码基因。为了提升美国国家生物信息中心(National Center for Biotechnology Information, NCBI)数据库中基因组的注释质量,本研究联合使用多种原核基因识别算法/软件和基因表达数据重注释1587个细菌和古细菌基因组。首先,利用Z曲线的33个变量从177个基因组原注释中识别获得3092个被过度注释为蛋白编码基因的序列;其次,通过同源比对为939个基因组中的4447个功能未知的蛋白编码基因注释上具体功能;最后,通过联合采用ZCURVE 3.0和Glimmer 3.02以及Prodigal这3种高精度的、广泛使用且基于算法不同而互补的基因识别软件来寻找漏注释基因。最终,从9个基因组中找到了2003个被漏注释的蛋白编码基因,这些基因属于多个蛋白质直系同源簇(clusters of orthologous groups of proteins, COG)。本研究使用新的工具并结合多组学数据重新注释早期测序的细菌和古细菌基因组,不仅为新测序菌株提供注释方法参考,而且这些重注释后得到的细菌基因序列也会对后续基础研究有所帮助。

本文引用格式

刘硕, 曾志, 曾凡才, 杜萌泽 . 基于序列相似性和Z曲线方法重注释原核生物蛋白编码基因[J]. 遗传, 2020 , 42(7) : 691 -702 . DOI: 10.16288/j.yczz.20-022

Abstract

The development of sequencing technology has generated huge genomic sequencing information and largely enriched public genetic resources. To analyze such big data, the algorithms and tools for comparison and annotation of genomes are updated continually, enabling genome annotation with higher accuracy via various annotation tools. Many prokaryotic genomes in public database were sequenced and assembled more than a decade ago, and they contained multiple genes with unknown functions. To improve the current annotation for those genomes in NCBI, we re-annotate 1587 bacterial and archaeal genomes using multiple prokaryotic gene recognition algorithms/softwares and gene expression data. The 33 Z-curve variables were applied to recognize sequences that were over-annotated to genes of 1587 bacterial and archaeal genomes deposited in public databases, and a total of 3092 sequences belonging to 177 genomes were recognized as sequences over-annotated as protein-coding genes. Next, 4447 protein-coding genes with unknown functions from 939 genomes were annotated with definite functions by similarity search. Finally, we recognized 2003 missed protein-coding genes that belong to known COG (clusters of orthologous groups of proteins) of nine genomes using three methods (ZCURVE 3.0, Glimmer 3.02 and Prodigal), which are accurate and frequently used for gene finding. Their algorithms are different and complementary. This is a comprehensive study for re-annotation of bacterial and archaeal genomes with new tools combining multi-omics data, which should provide a reference for annotation of newly sequenced strains, and also benefit further fundamental researches with the bacterial gene sequences obtained after re-annotation.

参考文献

[1] M?rk S, Holmes I . Evaluating bacterial gene-finding HMM structures as probabilistic logic programs. Bioinformatics, 2012,28(5):636-642.
[2] Warren AS, Archuleta J, Feng WC, Setubal JC . Missing genes in the annotation of prokaryotic genomes. BMC Bioinformatics, 2010,11(1):131.
[3] Salzberg SL . Next-generation genome annotation: we still struggle to get it right. Genome Biol, 2019,20(1):92.
[4] Breitwieser FP, Pertea M, Zimin AV, Salzberg SL . Human contamination in bacterial genomes has created thousands of spurious proteins. Genome Res, 2019,29(6):954-960.
[5] Fleischmann RD, Adams MD, White O, Clayton RA, Kirkness EF, Kerlavage AR, Bult CJ, Tomb JF, Dougherty BA, Merrick JM, Mckenney K, Sutton G, FitzHugh W,Fields C,Gocayne JD,Scott J,Shirley R,Liu LL,Glodek A,Kelley JM,Weidman JF,Phillips CA,Spriggs T,Hedblom E,Cotton MD,Utterback TR,Hanna MC,Nguyen DT,Saudek DM,Brandon RC,Fine LD,Fritchman JL,Fuhrmann JL,Geoghagen NSM,Gnehm CL,McDonald LA,Small KV,Fraser CM,Smith HO,Venter JC. Whole-genome random sequencing and assembly of Haemophilus influenzae Rd. Science, 1995,269(5223):496-512.
[6] Benson DA, Karsch-Mizrachi I, Lipman DJ, Ostell J, Sayers EW . Genbank. Nucleic Acids Res, 2009,37(Database issue):D26-D31.
[7] Yu JF, Xiao K, Jiang DK, Guo J, Wang JH, Sun X . An Integrative method for identifying the over-annotated protein-coding genes in microbial genomes. DNA Res, 2011,18(6):435-449.
[8] Hua ZG, Lin Y, Yuan YZ, Yang DC, Wei W, Guo FB . ZCURVE 3.0: identify prokaryotic genes with higher accuracy as well as automatically and accurately select essential genes. Nucleic Acids Res, 2015,43(W1):W85-W90.
[9] Zickmann F, Renard BY . IPred-integrating ab initio and evidence based gene predictions to improve prediction accuracy. BMC Genomics, 2015,16(1):134.
[10] Keilwagen J, Wenk M, Erickson JL, Schattat MH, Grau J, Hartung F . Using intron position conservation for homology-based gene prediction. Nucleic Acids Res, 2016,44(9):e89.
[11] Besemer J, Lomsadze A, Borodovsky M . GeneMarkS:a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions. Nucleic Acids Res, 2001,29(12):2607-2618.
[12] Kelley DR, Liu B, Delcher AL, Pop M, Salzberg SL . Gene prediction with Glimmer for metagenomic sequences augmented by classification and clustering. Nucleic Acids Res, 2012,40(1):e9.
[13] Larsen TS, Krogh A . EasyGene-a prokaryotic gene finder that ranks ORFs by statistical significance. BMC Bioinformatics, 2003,4(1):21.
[14] Guo FB, Ou HY, Zhang CT . ZCURVE: a new system for recognizing protein-coding genes in bacterial and archaeal genomes. Nucleic Acids Res, 2003,31(6):1780-1789.
[15] Du MZ, Guo FB, Chen YY . Gene re-annotation in genome of the extremophile Pyrobaculum Aerophilum by using bioinformatics methods. J Biomol Struct Dyn, 2011,29(2):391-401.
[16] Guo FB, Xiong LF, Teng JL, Yuen KY, Lau SK, Woo PC . Re-annotation of protein-coding genes in 10 complete genomes of Neisseriaceae family by combining similarity- based and composition-based methods. DNA Res, 2013,20(3):273-286.
[17] Lei Y, Kang SK, Gao JX, Jia XS, Chen LL . Improved annotation of a plant pathogen genome Xanthomonas oryzae pv. oryzae PXO99A. J Biomol Struct Dyn, 2013,31(3):342-350.
[18] Mao Y, Yang X, Liu Y, Yan Y, Du Z, Han Y, Song Y, Zhou L, Cui Y, Yang R . Reannotation of Yersinia pestis strain 91001 Based on Omics Data. Am J Trop Med Hyg, 2016,95(3):562-570.
[19] Pfeiffer F, Bagyan I, Alfaro‐Espinoza G,Zamora‐Lagos MA,Habermann B,Marin‐Sanguino A,Oesterhelt D,Kunte HJ. Revision and reannotation of the Halomonas elongata DSM 2581T genome. MicrobiologyOpen, 2017,6(4):e00465.
[20] Delcher AL, Bratke KA, Powers EC, Salzberg SL . Identifying bacterial genes and endosymbiont DNA with Glimmer. Bioinformatics, 2007,23(6):673-679.
[21] Zhang R, Zhang CT . A Brief Review:The Z-curve theory and its application in genome analysis. Curr Genomics, 2014,15(2):78-94.
[22] Weiss MC, Sousa FL, Mrnjavac N, Neukirchen S, Roettger M, Nelson-Sathi S, Martin WF . The physiology and habitat of the last universal common ancestor. Nat Microbiol, 2016,1(9):16116.
[23] Hyatt D, Chen GL, Locascio PF, Land ML, Larimer FW, Hauser LJ . Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics, 2010,11(1):119.
[24] Barrett TT, Troup DB, Wilhite SE, Ledoux P, Rudnev D, Evangelista C, Kim IF, Soboleva A, Tomashevsky M, Marshall KA, Phillippy KH, Sherman PM, Muertter RN, Edgar R . NCBI GEO: archive for high-throughput functional genomic data. Nucleic Acids Res, 2009,37(Database issue):D885-D890.
[25] Wang M, Weiss M, Simonovic M, Haertinger G, Schrimpf SP, Hengartner MO,von Mering C. PaxDb, a database of protein abundance averages across all three domains of life. Mol Cell Proteomics, 2012,11(8):492-500.
[26] McGinnis S, Madden TL . BLAST: at the core of a powerful and diverse set of sequence analysis tools. Nuleic Acids Res, 2004,32(Suppl.2):W20-W25.
[27] Wood DE, Lin H, Levy-Moonshine A, Swaminathan R, Chang YC, Anton BP, Osmani L, Steffen M, Kasif S, Salzberg SL . Thousands of missed genes found in bacterial genomes and their analysis with COMBREX. Biol Direct, 2012,7(1):37.
[28] Huerta-Cepas J, Szklarczyk D, Forslund K, Cook H, Heller D, Walker MC, Rattei T, Mende DR, Sunagawa S, Kuhn M, Jensen LJ, Mering CV, Bork P. eggNOG 4.5: a hierarchical orthology framework with improved functional annotations for eukaryotic, prokaryotic and viral sequences. Nucleic Acids Res, 2016,44(Database issue):D286-D293.
[29] Wu ST, Zhu ZW, Fu LM, Niu BF, Li WZ . WebMGA: a customizable web server for fast metagenomic sequence analysis. BMC Genomics, 2011,12(1):444.
[30] Tatusov RL, Galperin MY, Natale DA, Koonin EV . The COG database: a tool for genome-scale analysis of protein functions and evolution. Nucleic Acids Res, 2000,28(1):33-36.
[31] Qi J, Luo H, Hao BL . CVTree: a phylogenetic tree reconstruction tool based on whole genomes. Nucleic Acids Res, 2004,32:W45-W47.
[32] Hockenbery D, Nu?ez G, Milliman C, Schreiber RD, Korsmeyer SJ . Bcl-2 is an inner mitochondrial membrane protein that blocks programmed cell death. Nature, 1990,348(6299):334-336.
[33] Liu WQ, Feng Y, Wang Y, Zou QH, Chen F, Guo JT, Peng YH, Jin Y, Li YG, Hu SN, Johnson RN, Liu GR, Liu SL . Salmonella paratyphi C: Genetic divergence from Salmonella choleraesuis and pathogenic convergence with Salmonella typhi. PLoS One, 2009,4(2):e4510.
[34] Vankuren NW, Long M . Gene duplicates resolving sexual conflict rapidly evolved essential gametogenesis functions. Nat Ecol Evol, 2018,2(4):705-712.
[35] Minor LL, Bockemühl J . 1987 supplement (no.31) to the schema of Kauffmann-White. Ann Inst Pasteur Microbiol, 1988,139(3):331-335.
[36] Dwyer DJ, Belenky PA, Yang JH, Macdonald IC, Martell JD, Takahashi N, Chan CT,lobritz MA,Braff D,Schwarz EG,Ye JD,Pati M,Vercruysse M,Ralifo PS,Allison KR,Khalil AS,Ting AY,Walker GC,Collins JJ. Antibiotics induce redox-related physiological alterations as part of their lethality. Proc Natl Acad Sci USA, 2014,111(20):E2100-E2109.
[37] Hadjeras L, Poljak L, Bouvier M, Morin-Ogier Q, Canal l,Cocaign-Bousquet M,Girbal L,Carpousis AJ. Detachment of the RNA degradosome from the inner membrane of Escherichia coli results in a global slowdown of mRNA degradation, proteolysis of RNase E and increased turnover of ribosome-free transcripts. Mol Microbiol, 2019,111(6):1715-1731.
[38] Kim S, Yu Z, Kil RM, Lee M . Deep learning of support vector machines with class probability output networks. Neural Netw, 2015,64:19-28.
[39] Guo FB . The distribution patterns of bases of protein- coding genes, non-coding ORFs, and intergenic sequences in Pseudomonas aeruginosa PA01 genome and its implication. J Biomol Struct Dyn, 2007,25(2):127-133.
[40] Uyar B, Yusuf D, Wurmus R, Rajewsky N, Ohler U, Akalin A . RCAS: an RNA centric annotation system for transcriptome-wide regions of interest. Nucleic Acids Res, 2017,45(10):e91.
[41] Huang Y, Liu Q, Chi LJ, Shi CM, Wu Z, Hu M, Shi H, Chen H . Application of BIG-Annotator in the genome sequencing data functional annotation and genetic diagnosis. Hereditas(Beijing), 2018,40(11):1015-1023.
[41] 黄莹, 刘琪, 池连江, 石承民, 吴祯, 胡敏, 石宏, 陈华 . BIG-Annotator: 基因组测序数据高效功能注释及其在遗传诊断中的应用. 遗传, 2018,40(11):1015-1023.
[42] Bick JT, Zeng SQ, Robinson MD, Ulbrich SE, Bauersachs S . Mammalian Annotation Database for improved annotation and functional classification of Omics datasets from less well-annotated organisms. Database(Oxford), 2019,2019:1-16.
[43] Ravindran SP, Lüneburg J, Gottschlich L, Tams V, Cordellier M . Daphnia stressor database: Taking advantage of a decade of Daphnia ‘-omics’ data for gene annotation. Sci Rep, 2019,9(1):11135.
文章导航

/