测序深度和DP过滤对基于二代测序的SNP基因分型准确性的影响
闫锦坤和周文川并列第一作者。
收稿日期: 2026-03-20
修回日期: 2026-06-15
网络出版日期: 2026-07-07
基金资助
2026 年西安市农业关键技术攻关重点项目(2026-JH-NJZD-0002);2026 年陕西省农业农村厅农业科技创新项目(SNTX-13)
The impact of sequencing depth and DP filtering on genotyping accuracy of SNP based on next-generation sequencing
Received date: 2026-03-20
Revised date: 2026-06-15
Online published: 2026-07-07
Supported by
Xi’an Key Technology Research and Development Project for Agriculture(2026-JH-NJZD-0002);Agricultural Science and Technology Innovation Project of Shaanxi Provincial Department of Agriculture and Rural Affairs(SNTX-13)
二代测序(Next-generation sequencing,NGS)技术已成为大规模群体SNP基因分型的首要方法。然而,在基因组杂合度高、连锁不平衡程度低的群体中,平衡SNP基因分型准确性与测序成本之间的矛盾仍面临挑战。测序深度和覆盖深度(depth of coverage,DP)是影响SNP基因分型准确性的关键参数,但目前关于不同测序深度下SNP基因分型结果的一致性,以及DP过滤对低深度测序SNP基因分型准确性的影响仍较为缺乏。本研究以略阳乌鸡(n=11)为研究对象,从两个层面评估5×和10×测序深度SNP基因分型一致性:一是直接比较两种测序深度获得的SNP基因型,二是以SLCO1B3基因位点162个SNP的Sanger测序分型结果为金标准研究两种测序深度对SNP基因型检测的准确性。以10×基因型检测结果为参照,研究DP0~DP5对5×测序深度下SNP基因分型结果的影响。以染色体为单位,分析SNP分型结果与GC含量、重复序列的关系。结果显示,5×和10×的基因分型一致性为78.07%,5×与Sanger测序结果的一致性为78.40%,10×为87.17%。99.18%的不一致性结果出现在杂合子的分型上,其中5×AB-10×AA/BB亚型主要出现在14条长度低于10 Mb的小染色体上,且与重复序列、GC含量显著关联。随着DP阈值从0提升至5,一致性位点比例从78.07%提升至87.33%,但是SNP的检测率从100%减少至35.07%;缺失位点比例从13.16%下降至4.08%,不一致位点比例无显著变化。上述结果表明,二代测序的难点在于杂合子的检测上,分型的准确性受测序深度和染色体的序列特征协同影响。从物理层面提升测序深度能提高杂合子分型的准确性,而算法层面的提升作用不明显。提升DP值能显著提高低深度与高深度测序结果的一致性,但随DP值的升高提升作用快速衰减、SNP的检出率大幅减少。
闫锦坤, 周文川, 薛花明, 夏媛, 许亮, 王哲鹏 . 测序深度和DP过滤对基于二代测序的SNP基因分型准确性的影响[J]. 遗传, 2026 , 48(8) : 817 -827 . DOI: 10.16288/j.yczz.26-069
Next-generation sequencing (NGS) is the primary method for SNP genotyping in a large-scale population. However, it still faces a challenge to balance the paradox between SNP genotyping accuracy and sequencing cost in populations characterized by high heterozygosity and low genomic linkage disequilibrium. Although sequencing depth and depth of coverage (DP) are critical parameters influencing SNP genotyping accuracy, comprehensive investigation on genotyping consistency of SNP among different sequencing depths and impact of DP on SNP genotyping accuracy of low-depth sequencing is of absence. In this study, we utilized Lueyang Black-boned chickens (n=11) to evaluate genotyping consistency of SNP at 5× and 10× sequencing depths via two strategies: (1) comparison of SNP genotyping results between 5× and 10×; (2) comparison of genotyping accuracies of 5× and 10× by employing Sanger-genotyping results of 162 SNP at the SLCO1B3 locus as the criteria. Furthermore, we investigated the effect of DP filtering on genotyping consistency between 5× and 10× by increasing DP from 0 to 5. We studied the association of genotyping results of NGS with GC contents and repeat sequences for each chromosome. The results show that the ratio of genotyping consistency is 78.07% between 5× and 10×. The ratio is 78.40% between 5× and Sanger sequencing, and increased to 87.17% for 10×. Almost all (99.18%) of inconsistent results happen in heterozygotes, of which 5×AB-10×AA/BB is mainly present in 14 microchromosomes with lengths less than 10 Mb. The 5×AB-10×AA/BB subset of inconsistent results is significantly associated with repeat sequences and GC contents. The ratio of consistent loci increases from 78.07% to 87.33% with the increasing of DP from 0 to 5. However, calling rates of SNP dramatically reduces from 100% to 35.07%. The ratio of missing loci decreases from 13.16% to 4.08%. DP filtering has no significant effect on the ratio of inconsistent loci. The results indicate that challenge of the NGS-based SNP genotyping approach focuses on genotyping of heterozygotes. Sequencing depth and sequence features of chromosomes affect genotyping accuracy. To physically increase sequencing depth improves the genotyping accuracy of SNP, whereas to algorithmically raise DP has a negligible effect. Increasing DP can significantly improve the genotyping consistency between low- and high-depth sequencing. The improving effect quickly decays with the increasing of DP, whereas calling rates of SNP dramatically reduce.
Key words: next-generation sequencing; sequencing depth; depth of coverage; SNP; genotyping
| [1] | Kim S, Misra A. SNP genotyping: technologies and biomedical applications. Annu Rev Biomed Eng, 2007, 9: 289-320. |
| [2] | Dorado G, Gálvez S, Rosales TE, Vásquez VF, Hernández P. Analyzing modern biomolecules: the revolution of nucleic-acid sequencing - review. Biomolecules, 2021, 11(8): 1111. |
| [3] | Kan-Lingwood NY, Sagi L, Mazie S, Shahar N, Zecherle Bitton L, Templeton A, Rubenstein D, Bouskila A, Bar-David S. Genotyping error detection and customised filtration for SNP datasets. Mol Ecol Resour, 2025, 25(1): e14033. |
| [4] | Huang J, Liang XM, Xuan YK, Geng CY, Li YX, Lu HR, Qu SF, Mei XL, Chen HB, Yu T, Sun N, Rao JH, Wang JH, Zhang WW, Chen Y, Liao S, Jiang H, Liu X, Yang ZP, Mu F, Gao SX. A reference human genome dataset of the BGISEQ-500 sequencer. Gigascience, 2017, 6(5): 1-9. |
| [5] | Schmidt J, Berghaus S, Blessing F, Herbeck H, Blessing J, Schierack P, Rödiger S, Roggenbuck D, Wenzel F. Genotyping of familial mediterranean fever gene (MEFV)-single nucleotide polymorphism-comparison of nanopore with conventional Sanger sequencing. PLoS One, 2022, 17(3): e0265622. |
| [6] | Adams DR, Eng CM. Next-generation sequencing to diagnose suspected genetic disorders. N Engl J Med, 2018, 379(14): 1353-1362. |
| [7] | Wright CF, FitzPatrick DR, Firth HV. Paediatric genomics: diagnosing rare disease in children. Nat Rev Genet, 2018, 19(5): 253-268. |
| [8] | Adelson RP, Renton AE, Li WT, Barzilai N, Atzmon G, Goate AM, Davies P, Freudenberg-Hua Y. Empirical design of a variant quality control pipeline for whole genome sequencing data using replicate discordance. Sci Rep, 2019, 9(1): 16156. |
| [9] | Buermans HPJ, den Dunnen JT. Next generation sequencing technology: Advances and applications. Biochim Biophys Acta, 2014, 1842(10): 1932-1941. |
| [10] | Goodwin S, McPherson JD, McCombie WR. Coming of age: ten years of next-generation sequencing technologies. Nat Rev Genet, 2016, 17(6): 333-351. |
| [11] | Satam H, Joshi K, Mangrolia U, Waghoo S, Zaidi G, Rawool S, Thakare RP, Banday S, Mishra AK, Das G, Malonia SK. Next-generation sequencing technology: Current trends and advancements. Biology (Basel), 2023, 12(7): 997. |
| [12] | Das S, Abecasis GR, Browning BL. Genotype imputation from large reference panels. Annu Rev Genomics Hum Genet, 2018, 19: 73-96. |
| [13] | Phocas F. Genotyping, the usefulness of imputation to increase SNP density, and imputation methods and tools. Methods Mol Biol, 2022, 2467: 113-138. |
| [14] | Porcu E, Sanna S, Fuchsberger C, Fritsche LG. Genotype imputation in genome-wide association studies. Curr Protoc Hum Genet, 2013, Chapter 1: Unit 1.25. |
| [15] | Song K, Li L, Zhang GF. Coverage recommendation for genotyping analysis of highly heterologous species using next-generation sequencing technology. Sci Rep, 2016, 6: 35736. |
| [16] | Chiara M, Gioiosa S, Chillemi G, D'Antonio M, Flati T, Picardi E, Zambelli F, Horner DS, Pesole G, Castrignanò T. CoVaCS: a consensus variant calling system. BMC Genomics, 2018, 19(1): 120. |
| [17] | Liu J, Shen QM, Bao HG. Comparison of seven SNP calling pipelines for the next-generation sequencing data of chickens. PLoS One, 2022, 17(1): e0262574. |
| [18] | Gargis AS, Kalman L, Berry MW, Bick DP, Dimmock DP, Hambuch T, Lu F, Lyon E, Voelkerding KV, Zehnbauer BA, Agarwala R, Bennett SF, Chen B, Chin ELH, Compton JG, Das S, Farkas DH, Ferber MJ, Funke BH, Furtado MR, Ganova-Raeva LM, Geigenmüller U, Gunselman SJ, Hegde MR, Johnson PLF, Kasarskis A, Kulkarni S, Lenk T, Liu CSJ, Manion M, Manolio TA, Mardis ER, Merker JD, Rajeevan MS, Reese MG, Rehm HL, Simen BB, Yeakley JM, Zook JM, Lubin IM. Assuring the quality of next-generation sequencing in clinical laboratory practice. Nat Biotechnol, 2012, 30(11): 1033-1036. |
| [19] | Yu XQ, Sun SY. Comparing a few SNP calling algorithms using low-coverage sequencing data. BMC Bioinformatics, 2013, 14: 274. |
| [20] | Hwang S, Kim E, Lee I, Marcotte EM. Systematic comparison of variant calling pipelines using gold standard personal exome variants. Sci Rep, 2015, 5: 17875. |
| [21] | Luo RB, Schatz MC, Salzberg SL. 16GT: a fast and sensitive variant caller using a 16-genotype probabilistic model. Gigascience, 2017, 6(7): 1-4. |
| [22] | Wang ZP, Chen Q, Wang YW, Wang YL, Liu RF. Refine localizations of functional variants affecting eggshell color of Lueyang black-boned chicken in the SLCO1B3. Poult Sci, 2024, 103(1): 103212. |
| [23] | Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics, 2009, 25(14): 1754-1760. |
| [24] | Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, Whitwham A, Keane T, McCarthy SA, Davies RM, Li H. Twelve years of SAMtools and BCFtools. Gigascience, 2021, 10(2): giab008. |
| [25] | McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, Garimella K, Altshuler D, Gabriel S, Daly M, DePristo MA. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res, 2010, 20(9): 1297-1303. |
| [26] | Van der Auwera GA, Carneiro MO, Hartl C, Poplin R, Del Angel G, Levy-Moonshine A, Jordan T, Shakir K, Roazen D, Thibault J, Banks E, Garimella KV, Altshuler D, Gabriel S, DePristo MA. From FastQ data to high confidence variant calls: the Genome Analysis Toolkit best practices pipeline. Curr Protoc Bioinformatics, 2013, 43(1110): 11.10. 1-11.10.33. |
| [27] | Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MAR, Bender D, Maller J, Sklar P, de Bakker PIW, Daly MJ, Sham PC. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet, 2007, 81(3): 559-575. |
| [28] | Treangen TJ, Salzberg SL. Repetitive DNA and next-generation sequencing: computational challenges and solutions. Nat Rev Genet, 2011, 13(1): 36-46. |
| [29] | Chen YC, Liu T, Yu CH, Chiang TY, Hwang CC. Effects of GC bias in next-generation-sequencing data on de novo genome assembly. PLoS One, 2013, 8(4): e62856. |
| [30] | Delahaye C, Nicolas J. Sequencing DNA with nanopores: troubles and biases. PLoS One, 2021, 16(10): e0257521. |
| [31] | Huang Z, Xu ZX, Bai H, Huang YJ, Kang N, Ding XT, Liu J, Luo HR, Yang CT, Chen WJ, Guo QX, Xue LZ, Zhang XP, Xu L, Chen ML, Fu HG, Chen YL, Yue ZC, Fukagawa T, Liu SL, Chang GB, Xu LH. Evolutionary analysis of a complete chicken genome. Proc Natl Acad Sci USA, 2023, 120(8): e2216641120. |
| [32] | Carson AR, Smith EN, Matsui H, Brækkan SK, Jepsen K, Hansen JB, Frazer KA. Effective filtering strategies to improve data quality from population-based whole exome sequencing studies. BMC Bioinformatics, 2014, 15: 125. |
| [33] | Ott A, Liu SZ, Schnable JC, Yeh CTE, Wang KS, Schnable PS. tGBS® genotyping-by-sequencing enables reliable genotyping of heterozygous loci. Nucleic Acids Res, 2017, 45(21): e178. |
/
| 〈 |
|
〉 |