[an error occurred while processing this directive]
Technique and Method

Automatic analysis pipeline of next-generation sequencing data

Expand
  • 1. State Key Laboratory of Cardiovascular Disease, Fuwai Hospital, National Center for Cardiovascular Disease, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing 100037, China;
    2. College of Biomedical Engineering, South-Central University for Nationalities, Wuhan 430074, China

Received date: 2013-09-07

  Revised date: 2014-01-20

  Online published: 2014-05-28

Abstract

The development of next-generation sequencing has generated high demand for data processing and analysis. Although there are a lot of software for analyzing next-generation sequencing data, most of them are designed for one specific function (e.g., alignment, variant calling or annotation). Therefore, it is necessary to combine them together for data analysis and to generate interpretable results for biologists. This study designed a pipeline to process Illumina sequencing data based on Perl programming language and SGE system. The pipeline takes original sequence data (fastq format) as input, calls the standard data processing software (e.g., BWA, Samtools, GATK, and Annovar), and finally outputs a list of annotated variants that researchers can further analyze. The pipeline simplifies the manual operation and improves the efficiency by automatization and parallel computation. Users can easily run the pipeline by editing the configuration file or clicking the graphical interface. Our work will facilitate the research projects using the sequencing technology.

Cite this article

Wenke Li, Fengyu Li, Siyao Zhang, Bin Cai, Na Zheng, Yu Nie, Dao Zhou, Qian Zhao . Automatic analysis pipeline of next-generation sequencing data[J]. Hereditas(Beijing), 2014 , 36(6) : 618 -624 . DOI: 10.3724/SP.J.1005.2014.0618

References

[1]Illumina Inc. Illumina Sequencing Technology. http://www. illumina.com/documents/products/techspotlights/techspotlight_sequencing.pdf.
[2]Cock PJA, Fields CJ, Goto N, Heuer ML, Rice PM. The Sanger FASTQ file format for sequences with quality scores, and the Solexa/Illumina FASTQ variants. Nucleic Acids Res, 2010, 38(6): 1767-1771.
[3]Li H, Durbin R. Fast and accurate long-read alignment with Burrows-Wheeler Transform. Bioinformatics, 2010, 26(5): 589-595.
[4]Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer
[5]N, Marth G, Abecasis G, Durbin R, 1000 Genome Project Data Processing Subgroup. The sequence alignment/map (SAM) format and SAMtools. Bioinformatics, 2009, 25(16): 2078-2079.
[6]Picard. http://picard.sourceforge.net
[7]McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, Garimella K, Altshuler D, Gabriel S, Daly M, DePristo MA. The genome analysis toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res, 2010, 20(9): 1297-1303.
[8]Wang K, Li M, Hakonarson H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res, 2010, 38(16): e164.
[9]Scbwartz RL, Pboenix T, Foy BD著. 盛春, 蒋永清, 王晖译. Perl语言入门 (第五版). 南京: 东南大学出版社, 2009, 200.
[10]ORACLE INC. N1 Grid Engine 6 用户指南. http://docs.oracle.com/cd/E19080-01/n1.grid.eng6/817-7681/esqcr/index.html.

Outlines

/