基于单核苷酸多态性的基因互作分析方法学进展
收稿日期: 2013-06-08
修回日期: 2013-07-24
网络出版日期: 2013-11-25
基金资助
国家自然科学基金项目(编号:30830104, 31071166), 广东省科技计划攻关项目(编号:2009A030301004), 东莞市科技重点项目(编号:201108101015)和广东医学院基金项目(编号:XG1001, XZ1105, STIF201122)资助
Advances in development of gene-gene interaction analy-sis methods based on SNP data: a review
Received date: 2013-06-08
Revised date: 2013-07-24
Online published: 2013-11-25
基于单核苷酸多态性的关联分析已成为当前解析人类常见复杂疾病遗传机制的重要手段之一, 然而, 目前普遍使用的单位点分析策略仅能发现部分单独效应显著的易感SNP位点, 因此遗漏了重要的遗传力组分——基因上位效应或联合效应。识别全基因组多基因间复杂的互作关系已成为全面解析复杂疾病致病分子机制必不可少的一项任务。已有很多方法被应用于全基因组交互作用分析, 加深了人类对复杂疾病遗传机制的进一步认识。基于各类方法的理论基础及算法的异同, 文章对目前应用较为广泛的基于遗传互作模型的方法、不基于互作模型的方法和数据挖掘类算法3类方法进行了系统地评述, 着重介绍了这些方法的主要思想、实现过程及应用中的注意事项等, 并指出开展大规模全基因组范围互作检测面临的问题, 以期能为相关领域的研究者提供方法学参考。
栾奕昭 左晓宇 刘轲 李谷 饶绍奇 . 基于单核苷酸多态性的基因互作分析方法学进展[J]. 遗传, 2013 , 35(12) : 1331 -1339 . DOI: 10.3724/SP.J.1005.2013.01331
The SNP-based association analysis has become one of the most important approaches to interpret the underlying molecular mechanisms for human complex diseases. Nevertheless, the widely-used singe-locus analysis is only capable of capturing a small portion of susceptible SNPs with prominent marginal effects, leaving the important genetic component, epistasis or joint effects, to be undetectable. Identifying the complex interplays among multiple genes in the genome-wide context is an essential task for systematically unraveling the molecular mechanisms for complex diseases. Many approaches have been used to detect genome-wide gene-gene interactions and provided new insights into the genetic basis of complex diseases. This paper reviewed recent advances of the methods for detecting gene-gene interaction, categorized into three types, model-based and model-free statistical methods, and data mining methods, based on their characteristics in theory and numerical algorithm. In particular, the basic principle, numerical implementation and cautions for application for each method were elucidated. In addition, this paper briefly discussed the limitations and challenges associated with detecting genome-wide epistasis, in order to provide some methodological consultancies for scientists in the related fields.
[1] Cordell HJ. Detecting gene-gene interactions that underlie human diseases. Nat Rev Genet, 2009, 10(6): 392–404.
[2] Moore JH. The ubiquitous nature of epistasis in determining susceptibility to common human diseases. Hum Hered, 2003, 56(1–3): 73–82.
[3] Marchini J, Donnelly P, Cardon LR. Genome-wide strategies for detecting multiple loci that influence complex diseases. Nat Genet, 2005, 37(4): 413–417.
[4] Briggs FBS, Ramsay PP, Madden E, Norris JM, Holers VM, Mikuls TR, Sokka T, Seldin MF, Gregersen PK, Criswell LA, Barcellos LF. Supervised machine learning and logistic regression identifies novel epistatic risk factors with PTPN22 for rheumatoid arthritis. Genes Immun, 2010, 11(3): 199–208.
[5] Epstein MP, Satten GA. Inference on haplotype effects in case-control studies using unphased genotype data. Am J Hum Genet, 2003, 73(6): 1316–1329.
[6] Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MA, Bender D, Maller J, Sklar P, de Bakker PI, Daly MJ, Sham PC. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet, 2007, 81(3): 559–575.
[7] Wu XS, Dong H, Luo L, Zhu Y, Peng G, Reveille JD, Xiong MM. A novel statistic for genome-wide interaction analysis. PLoS Genet, 2010, 6(9): e1001131.
[8] Ueki M, Cordell HJ. Improved statistics for genome-wide interaction analysis. PLoS Genet, 2012, 8(4): e1002625.
[9] Rao SQ, Yuan MQ, Zuo XY, Su WY, Zhang F, Huang K, Lin MH, Ding YL. A novel evolution-based method for detecting gene-gene interactions. PLoS ONE, 2011, 6(10): e26435.
[10] Breiman L. Random forests. Machine Learning, 2001, 45(1): 5–32.
[11] Bureau A, Dupuis J, Falls K, Lunetta KL, Hayward B, Keith TP, Van Eerdewegh P. Identifying SNPs predictive of phenotype using random forests. Genet Epidemiol, 2005, 28(2): 171–182.
[12] Cook N, Zee R, Ridker P. Tree and spline based association analysis of gene-gene interaction models for ischemic stroke. Stat Med, 2004, 23(9): 1439–1453.
[13] Lunetta K, Hayward L, Segal J, Van Eerdewegh P. Screening large-scale association study data exploiting interactions using random forests. BMC Genet, 2004, 5: 32.
[14] Ritchie MD, Hahn LW, Roodi N, Bailey LR, Dupont WD, Parl FF, Moore JH. Multifactor-dimensionality reduction reveals high-order interactions among estrogen-metabolism genes in sporadic breast cancer. Am J Hum Genet, 2001, 69(1): 138–147.
[15] Chung YJ, Lee SY, Elston RC, Park T. Odds ratio based multifactor-dimensionality reduction method for detecting gene-gene interactions. Bioinformatics, 2007, 23(1): 71–76.
[16] Collins RL, Hu T, Wejse C, Sirugo G, Williams SM, Moore JH. Multifactor dimensionality reduction reveals a three- locus epistatic interaction associated with susceptibility to pulmonary tuberculosis. BioData Min, 2013, 6(1): 4.
[17] Nunkesser R, Bernholt T, Schwender H, Ickstadt K, Wegener I. Detecting high-order interactions of single nucleotide polymorphisms using genetic programming. Bioinformatics, 2007, 23(24): 3280–3288.
[18] Liu KH, Xu CG. A genetic programming-based approach to the classification of multiclass microarray datasets. Bioinformatics, 2009, 25(3): 331–337.
[19] Ritchie MD, White BC, Parker JS, Hahn LW, Moore JH. Optimization of neural network architecture using genetic programming improves detection and modeling of gene-gene interactions in studies of human diseases. BMC Bioinformatics, 2003, 4(1): 28.
[20] Chen SH, Sun JL, Dimitrov L, Turner AR, Adams TS, Meyers DA, Chang BL, Zheng SL, Gronberg H, Xu JF, Hsu FC. A support vector machine approach for detecting gene-gene interaction. Genet Epidemiol, 2008, 32(2): 152– 167.
[21] McKinney BA, Reif D, Ritchie M, Moore JH. Machine learning for detecting gene-gene interactions. Appl Bioinformatics, 2006, 5(2): 77
/
| 〈 |
|
〉 |