[an error occurred while processing this directive]
en

A new alignment-free sequence analysis based on the distribution of K-tuple

Expand
  • 1. College of Science, Northwest A&F University, Yangling 712100, China;
     2. College of Life Science, Northwest A&F University, Yangling 712100, China

Received date: 2009-09-15

  Revised date: 2009-12-02

  Online published: 2010-05-24

Abstract

Based on the distribution of K-tuple in complete genome, a method without doing sequence alignment to infer difference of biological sequence is proposed in this paper. The method can be used to measure the difference of distribution on K-tuple between the native DNA sequences and the corresponding randomized ones. Applied to construct phylogenetic trees of the complete mitochondrial genomes of 26 species of placental mammals, with K increasing, it yields phylogenetic trees of which the classification effect increasingly matches the result widely recognised by the biological field. The final results show that the phylogenetic trees built by this method is more reasonable than by other alignment-free sequence methods.

Cite this article

CHEN Juan, TUN Wen-Wu, JIE Xiao-Chi, GUO Man-Cai, YUAN Zhi-Fa . A new alignment-free sequence analysis based on the distribution of K-tuple[J]. Hereditas(Beijing), 2010 , 32(6) : 606 -612 . DOI: 10.3724/SP.J.1005.2010.00606

References

[1] Qi J, Wang B, Hao BL. Whole proteome prokaryote phy-logeny without sequence alignment: A K-string composi-tion approach. J Mol Evol, 2004, 58(1): 1-11.

[2] Yu ZG, Jiang P. Distance, correlation and mutual informa-tion among portraits of organisms based on complete ge-nomes. Physics Lett A, 2001, 286(1): 34-46.

[3] 罗辽复, 纪丰民, 谢立青, 李弘谦. 以核苷酸短程关联为基础的进化树重建. 内蒙古大学学报(自然科学版), 2001, 32(1): 32-41.

[4] Wu TJ, BurKe JP, Davison DB. A measure of DNA se-quence dissimilarity based on Mahalanobis distance be-tween frequencies of words. Biometrics, 1997, 53(4): 1431-1439.

[5] Kull B, Leibler RA. On information and sufficiency. Ann Math Statist, 1951, 22(1): 79-86.

[6] 沈世镒, 陈鲁生. 信息论与编码理论. 北京: 北京科学出版社, 2002, 31-36.

[7] 刘军, 许甫荣. 基于相对熵原理构建生物进化系统树. 北京大学学报(自然科学版), 2003, 39(增刊): 76-80.

[8] Stormo GD. DNA binding sites: Representation and dis-covery. Bioinformatics, 2000, 16(1): 16-23.

[9] Li M, Badger JH, Chen X, Kwong S, Kearney P, Zhang HY. An information-based sequence distance and its ap-plication to whole mitochondrial genome phylogeny. Bioinformatics, 2001, 17(2): 149-154.

[10] 傅强, 钱敏平, 陈良标, 朱玉贤. 编码序列和非编码序列的3-tuple分布特征. 遗传学报, 2005, 32(10): 1018-1026.

[11] 李蒨, 李逢博, 王炜. 蛋白质序列复杂性简化与非比对序列分析. 生物化学与生物物理进展, 2006, 33(12): 1215-1222.

[12] Xie HM, Hao BL. Visualiztion of K-tuple distribution in pro-caryote complete genomes and their randomized counterparts. Proc IEEE Comput Soc Bioinform Conf, 2002, 1: 31-42.

[13] Hao BL, Xie HM, Yu ZG, Chen GY. Avoided string in bacterial complete genomes and a related combinatorial problem. Ann Combinatorics, 2000, 4(3-4): 247-255.

[14] 罗辽复. 生命进化的物理观. 上海: 上海科学技术出版社, 2000. 168-200.

[15] 贾晓超, 李培芳, 罗辽复. 基因组中“K字”频数的分布. 内蒙古大学学报(自然科学版), 2005, 36(3): 301-395.

[16] 罗辽复. DNA信息内容的普适关系. 合肥学院学报(自然科学版), 2005, 15(1): 1-6.

[17] Wu CF. The distribution of the frequency of occurrence of nucleotide subsequences. Methodol Comput Appl Probabil, 2005, 7: 325-334.

[18] Cao Y, Janke A, Waddell PJ, Westerman M, Takenaka O, Murata S, Okada N, Pääbo S, Hasegawa M. Conflict among individual mitochondrial proteins in resolving the phylogeny of Eutherian orders. J Mol Evol, 1998, 47(3): 307-322.

[19] Reyes A, Gissi C, Pesole G, Catzeflis FM, Saccone C. Where do rodents fit? Evidence from the complete mito-chondrial genome of Sciurus vulgaris. Mol Biol Evol, 2000, 17(6): 979-983.

[20] Li B, Li YB, He HB. LZ complexity distance of DNA se-quences and its application in phylogenetic tree reconstruction. Genomics Proteomics Bioinformatics, 2005, 3(4): 206-212.

[21] 李斌, 李义兵, 何红波. 符号序列间的LZ复杂性距离及其应用. 小型微型计算机系统, 2007, 28(5): 850-854.

Outlines

/