技术与方法

用于高通量测序的基因组靶序列捕获方法的建立

展开
  • 1. 上海交通大学医学院附属瑞金医院, 医学基因组学国家重点实验室, 上海 200025; 
    2. 生物芯片上海国家工程研究中心, 上海 201203; 
    3. 国家人类基因组南方研究中心, 上海 201203; 
    4. 上海交通大学医学院基础医学院病理教研室, 上海 200025

收稿日期: 2010-04-23

  修回日期: 2010-09-30

  网络出版日期: 2010-12-20

基金资助

国家重大科学研究项目(2006CB910402)资助

Establishment of target genomic DNA capturing system for next generation sequencing

Expand
  • 1. State Key Laboratory of Medical Genomics, Ruijin Hospital, Shanghai Jiaotong University School of Medicine, Shanghai 200025, China; 
    2. National Engineering Research Center for Biochip at Shanghai, Shanghai 201203, China; 
    3. Chinese National Human Genome Center at Shanghai, Shanghai 201203, China; 
    4. Department of Pathology, College of Basic Medical Science, Shanghai Jiaotong University School of Medicine, Shanghai 200025, China

Received date: 2010-04-23

  Revised date: 2010-09-30

  Online published: 2010-12-20

摘要

文章旨在建立一种基因组目标靶序列捕捉文库的方法, 并结合第二代测序技术, 以实现候选基因区段的深度测序。利用Agilent公司的eArray在线平台, 对1 250个基因的11 824个外显子共2 414 977 bp的基因组序列进行120个碱基长度的捕捉探针(钓饵)设计, 并制备成SureSelect液相靶序列捕获试剂。选用2例人基因组DNA, 超声打断后末端补平并磷酸化, 连接SOLiD接头, 回收150bp~200bp的DNA片段, 与靶序列探针杂交捕获目标序列, 油包水微乳滴PCR扩增后, 磁珠分离富集, 上SOLiD测序系统通过工作流程分析(WFA)进行文库质量的评价, 或正式测序反应。结果显示对所包含的11 147个基因外显子片段设计出并合成了46 509个捕捉探针, 制备成SureSelect试剂盒。探针可有效地捕捉并富集基因组DNA的目标靶片段, 定量PCR显示富集效率可达29倍。WFA分析表明文库可以在SOLiD仪器进行正式测序。测序结果显示靶序列区域的测序数占有效总测序数的比例达到70%, 覆盖率均在200×以上。结果表明本研究所建立的SureSelect基因组靶序列捕捉、富集建立测序文库的技术路线可行, 可直接用于SOLiD测序仪的测序。

本文引用格式

陈丹,张雯,朱智东,黄银,王平,周贝贝,杨晓楠,肖华胜,张庆华 . 用于高通量测序的基因组靶序列捕获方法的建立[J]. 遗传, 2010 , 32(12) : 1296 -1303 . DOI: 10.3724/SP.J.1005.2010.01296

Abstract

The motivation of this research is to establish a system of target genomic DNA capture and enrichment, which could be used in deep sequencing of target regions with next-generation sequencing. To design the 120 bp capture probes (baits) and prepare the SureSelect reagents, 2 414 977 bp human genomic sequence of 11 824 exons in 1 250 genes were submitted to the Agilent eArray platform and manufactured by Agilent. Two human genomic DNA samples were used and conducted the successive experiments for sequencing library construction: shearing fragmentation by sonication, blunt-ending and phosphorylation, adaptor ligation, 150?200 bp fragments size selection, followed by hybridization with the baits, hybrid selection with magnetic beads, and PCR amplification. Prior to SOLiD sequencing reaction, the libraries were amplified with emulsion PCR and enriched with the P2 enrichment beads. The library samples were loaded to sequencing Chip for Work Flow Analysis (WFA) or sequencing running with default parameters. The results displayed that 46 509 baits were designed and synthesized for 11 147 gene regions, and SureSelect capture probe regent was prepared. Real-time PCR showed the target enrichment efficiency up to 29 times with the SureSelect system. WFA revealed that the libraries were suitable for SOLiD Sequencing. The sequencing data revealed that 70% of the unique mapped sequence tags matched the target regions, and the average coverage of the target regions were above 200-fold. All these demonstrated the feasibility of the established system of target genome sequence capture for next generation DNA sequencing.

参考文献

[1] Antipova AA, Sokolsky TD, Clouser CR, Dimalanta ET, Hendrickson CL, Kosnopo C, Lee CC, Ranade SS, Zhang L, Blanchard AP, McKernan KJ. Polymorphism discovery in high-throughput resequenced microarray-enriched human genomic loci. J Biomol Tech, 2009, 20(5): 253–257. [2] Marth GT, Korf I, Yandell MD, Yeh RT, Gu Z, Zakeri H, Stitziel NO, Hillier L, Kwok PY, Gish WR. A general approach to single-nucleotide polymorphism discovery. Nat Genet, 1999, 23(4): 452–456. [3] Zhang J, Xiao L, Yin YF, Sirois P, Gao HL, Li K. A law of mutation: power decay of small insertions and small deletions associated with human diseases. Appl Biochem Biotechnol, 2010, 162(2): 321–328. [4] Kamb A. Mutation load, functional overlap, and synthetic lethality in the evolution and treatment of cancer. J Theor Biol, 2003, 223(2): 205–213. [5] Sachidanandam R, Weissman D, Schmidt SC, Kakol JM, Stein LD, Marth G, Sherry S, Mullikin JC, Mortimore BJ, Willey DL, Hunt SE, Cole CG, Coggill PC, Rice CM, Ning Z, Rogers J, Bentley DR, Kwok PY, Mardis ER, Yeh RT, Schultz B, Cook L, Davenport R, Dante M, Fulton L, Hillier L, Waterston RH, McPherson JD, Gilman B, Schaffner S, Van Etten WJ, Reich D, Higgins J, Daly MJ, Blumenstiel B, Baldwin J, Stange-Thomann N, Zody MC, Linton L, Lander ES, Altshuler D. A map of human genome sequence variation containing 1.42 million single nucleotide polymorphisms. Nature, 2001, 409(6822): 928–933. [6] Iafrate AJ, Feuk L, Rivera MN, Listewnik ML, Donahoe PK, Qi Y, Scherer SW, Lee C. Detection of large-scale variation in the human genome. Nat Genet, 2004, 36(9): 949–951. [7] Kidd JM, Cooper GM, Donahue WF, Hayden HS, Sampas N, Graves T, Hansen N, Teague B, Alkan C, Antonacci F, Haugen E, Zerr T, Yamada NA, Tsang P, Newman TL, Tüzün E, Cheng Z, Ebling HM, Tusneem N, David R, Gillett W, Phelps KA, Weaver M, Saranga D, Brand A, Tao W, Gustafson E, McKernan K, Chen L, Malig M, Smith JD, Korn JM, McCarroll SA, Altshuler DA, Peiffer DA, Dorschner M, Stamatoyannopoulos J, Schwartz D, Nickerson DA, Mullikin JC, Wilson RK, Bruhn L, Olson MV, Kaul R, Smith DR, Eichler EE. Mapping and sequencing of structural variation from eight human genomes. Nature, 2008, 453(7191): 56–64. [8] 何永蜀, 张闻, 杨照青. 人类基因组结构变异. 遗传, 2009, 31(8): 771–778. [9] Tuzun E, Sharp AJ, Bailey JA, Kaul R, Morrison VA, Pertz LM, Haugen E, Hayden H, Albertson D, Pinkel D, Olson MV, Eichler EE. Fine-scale structural variation of the human genome. Nat Genet, 2005, 37(7): 727–732. [10] Redon R, Ishikawa S, Fitch KR, Feuk L, Perry GH, Andrews TD, Fiegler H, Shapero MH, Carson AR, Chen W, Cho EK, Dallaire S, Freeman JL, González JR, Gratacòs M, Huang J, Kalaitzopoulos D, Komura D, MacDonald JR, Marshall CR, Mei R, Montgomery L, Nishimura K, Okamura K, Shen F, Somerville MJ, Tchinda J, Valsesia A, Woodwark C, Yang FT, Zhang JJ, Zerjal T, Zhang J, Armengol L, Conrad DF, Estivill X, Tyler-Smith C, Carter NP, Aburatani H, Lee C, Jones KW, Scherer SW, Hurles ME. Global variation in copy number in the human genome. Nature, 2006, 444(7118): 444–454. [11] Wong KK, deLeeuw RJ, Dosanjh NS, Kimm LR, Cheng Z, Horsman DE, MacAulay C, Ng RT, Brown CJ, Eichler EE, Lam WL. A comprehensive analysis of common copy- number variations in the human genome. Am J Hum Genet, 2007, 80(1): 91–104. [12] Rosa-Rosa JM, Gracia-Aznarez FJ, Hodges E, Pita G, Rooks M, Xuan Z, Bhattacharjee A, Brizuela L, Silva JM, Hannon GJ, Benitez J. Deep sequencing of target linkage assay-identified regions in familial breast cancer: methods, analysis pipeline and troubleshooting. PLoS One, 2010, 5(4): e9976. [13] Ng SB, Turner EH, Robertson PD, Flygare SD, Bigham AW, Lee C, Shaffer T, Wong M, Bhattacharjee A, Eichler EE, Bamshad M, Nickerson DA, Shendure J. Targeted capture and massively parallel sequencing of 12 human exomes. Natu
文章导航

/