[an error occurred while processing this directive]
Research Article

The impact of sequencing depth and DP filtering on genotyping accuracy of SNP based on next-generation sequencing

Expand
  • 1 College of Animal Science and Technology, Northwest A&F University, Yangling 712100, China
    2 Quality Safety Testing Center of Agricultural Product of Weinan City, Weinan 714000, China

Received date: 2026-03-20

  Revised date: 2026-06-15

  Online published: 2026-07-07

Supported by

Xi’an Key Technology Research and Development Project for Agriculture(2026-JH-NJZD-0002);Agricultural Science and Technology Innovation Project of Shaanxi Provincial Department of Agriculture and Rural Affairs(SNTX-13)

Abstract

Next-generation sequencing (NGS) is the primary method for SNP genotyping in a large-scale population. However, it still faces a challenge to balance the paradox between SNP genotyping accuracy and sequencing cost in populations characterized by high heterozygosity and low genomic linkage disequilibrium. Although sequencing depth and depth of coverage (DP) are critical parameters influencing SNP genotyping accuracy, comprehensive investigation on genotyping consistency of SNP among different sequencing depths and impact of DP on SNP genotyping accuracy of low-depth sequencing is of absence. In this study, we utilized Lueyang Black-boned chickens (n=11) to evaluate genotyping consistency of SNP at 5× and 10× sequencing depths via two strategies: (1) comparison of SNP genotyping results between 5× and 10×; (2) comparison of genotyping accuracies of 5× and 10× by employing Sanger-genotyping results of 162 SNP at the SLCO1B3 locus as the criteria. Furthermore, we investigated the effect of DP filtering on genotyping consistency between 5× and 10× by increasing DP from 0 to 5. We studied the association of genotyping results of NGS with GC contents and repeat sequences for each chromosome. The results show that the ratio of genotyping consistency is 78.07% between 5× and 10×. The ratio is 78.40% between 5× and Sanger sequencing, and increased to 87.17% for 10×. Almost all (99.18%) of inconsistent results happen in heterozygotes, of which 5×AB-10×AA/BB is mainly present in 14 microchromosomes with lengths less than 10 Mb. The 5×AB-10×AA/BB subset of inconsistent results is significantly associated with repeat sequences and GC contents. The ratio of consistent loci increases from 78.07% to 87.33% with the increasing of DP from 0 to 5. However, calling rates of SNP dramatically reduces from 100% to 35.07%. The ratio of missing loci decreases from 13.16% to 4.08%. DP filtering has no significant effect on the ratio of inconsistent loci. The results indicate that challenge of the NGS-based SNP genotyping approach focuses on genotyping of heterozygotes. Sequencing depth and sequence features of chromosomes affect genotyping accuracy. To physically increase sequencing depth improves the genotyping accuracy of SNP, whereas to algorithmically raise DP has a negligible effect. Increasing DP can significantly improve the genotyping consistency between low- and high-depth sequencing. The improving effect quickly decays with the increasing of DP, whereas calling rates of SNP dramatically reduce.

Cite this article

Jinkun Yan, Wenchuan Zhou, Huaming Xue, Yuan Xia, Liang Xu, Zhepeng Wang . The impact of sequencing depth and DP filtering on genotyping accuracy of SNP based on next-generation sequencing[J]. Hereditas(Beijing), 2026 , 48(8) : 817 -827 . DOI: 10.16288/j.yczz.26-069

References

[1] Kim S, Misra A. SNP genotyping: technologies and biomedical applications. Annu Rev Biomed Eng, 2007, 9: 289-320.
[2] Dorado G, Gálvez S, Rosales TE, Vásquez VF, Hernández P. Analyzing modern biomolecules: the revolution of nucleic-acid sequencing - review. Biomolecules, 2021, 11(8): 1111.
[3] Kan-Lingwood NY, Sagi L, Mazie S, Shahar N, Zecherle Bitton L, Templeton A, Rubenstein D, Bouskila A, Bar-David S. Genotyping error detection and customised filtration for SNP datasets. Mol Ecol Resour, 2025, 25(1): e14033.
[4] Huang J, Liang XM, Xuan YK, Geng CY, Li YX, Lu HR, Qu SF, Mei XL, Chen HB, Yu T, Sun N, Rao JH, Wang JH, Zhang WW, Chen Y, Liao S, Jiang H, Liu X, Yang ZP, Mu F, Gao SX. A reference human genome dataset of the BGISEQ-500 sequencer. Gigascience, 2017, 6(5): 1-9.
[5] Schmidt J, Berghaus S, Blessing F, Herbeck H, Blessing J, Schierack P, Rödiger S, Roggenbuck D, Wenzel F. Genotyping of familial mediterranean fever gene (MEFV)-single nucleotide polymorphism-comparison of nanopore with conventional Sanger sequencing. PLoS One, 2022, 17(3): e0265622.
[6] Adams DR, Eng CM. Next-generation sequencing to diagnose suspected genetic disorders. N Engl J Med, 2018, 379(14): 1353-1362.
[7] Wright CF, FitzPatrick DR, Firth HV. Paediatric genomics: diagnosing rare disease in children. Nat Rev Genet, 2018, 19(5): 253-268.
[8] Adelson RP, Renton AE, Li WT, Barzilai N, Atzmon G, Goate AM, Davies P, Freudenberg-Hua Y. Empirical design of a variant quality control pipeline for whole genome sequencing data using replicate discordance. Sci Rep, 2019, 9(1): 16156.
[9] Buermans HPJ, den Dunnen JT. Next generation sequencing technology: Advances and applications. Biochim Biophys Acta, 2014, 1842(10): 1932-1941.
[10] Goodwin S, McPherson JD, McCombie WR. Coming of age: ten years of next-generation sequencing technologies. Nat Rev Genet, 2016, 17(6): 333-351.
[11] Satam H, Joshi K, Mangrolia U, Waghoo S, Zaidi G, Rawool S, Thakare RP, Banday S, Mishra AK, Das G, Malonia SK. Next-generation sequencing technology: Current trends and advancements. Biology (Basel), 2023, 12(7): 997.
[12] Das S, Abecasis GR, Browning BL. Genotype imputation from large reference panels. Annu Rev Genomics Hum Genet, 2018, 19: 73-96.
[13] Phocas F. Genotyping, the usefulness of imputation to increase SNP density, and imputation methods and tools. Methods Mol Biol, 2022, 2467: 113-138.
[14] Porcu E, Sanna S, Fuchsberger C, Fritsche LG. Genotype imputation in genome-wide association studies. Curr Protoc Hum Genet, 2013, Chapter 1: Unit 1.25.
[15] Song K, Li L, Zhang GF. Coverage recommendation for genotyping analysis of highly heterologous species using next-generation sequencing technology. Sci Rep, 2016, 6: 35736.
[16] Chiara M, Gioiosa S, Chillemi G, D'Antonio M, Flati T, Picardi E, Zambelli F, Horner DS, Pesole G, Castrignanò T. CoVaCS: a consensus variant calling system. BMC Genomics, 2018, 19(1): 120.
[17] Liu J, Shen QM, Bao HG. Comparison of seven SNP calling pipelines for the next-generation sequencing data of chickens. PLoS One, 2022, 17(1): e0262574.
[18] Gargis AS, Kalman L, Berry MW, Bick DP, Dimmock DP, Hambuch T, Lu F, Lyon E, Voelkerding KV, Zehnbauer BA, Agarwala R, Bennett SF, Chen B, Chin ELH, Compton JG, Das S, Farkas DH, Ferber MJ, Funke BH, Furtado MR, Ganova-Raeva LM, Geigenmüller U, Gunselman SJ, Hegde MR, Johnson PLF, Kasarskis A, Kulkarni S, Lenk T, Liu CSJ, Manion M, Manolio TA, Mardis ER, Merker JD, Rajeevan MS, Reese MG, Rehm HL, Simen BB, Yeakley JM, Zook JM, Lubin IM. Assuring the quality of next-generation sequencing in clinical laboratory practice. Nat Biotechnol, 2012, 30(11): 1033-1036.
[19] Yu XQ, Sun SY. Comparing a few SNP calling algorithms using low-coverage sequencing data. BMC Bioinformatics, 2013, 14: 274.
[20] Hwang S, Kim E, Lee I, Marcotte EM. Systematic comparison of variant calling pipelines using gold standard personal exome variants. Sci Rep, 2015, 5: 17875.
[21] Luo RB, Schatz MC, Salzberg SL. 16GT: a fast and sensitive variant caller using a 16-genotype probabilistic model. Gigascience, 2017, 6(7): 1-4.
[22] Wang ZP, Chen Q, Wang YW, Wang YL, Liu RF. Refine localizations of functional variants affecting eggshell color of Lueyang black-boned chicken in the SLCO1B3. Poult Sci, 2024, 103(1): 103212.
[23] Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics, 2009, 25(14): 1754-1760.
[24] Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, Whitwham A, Keane T, McCarthy SA, Davies RM, Li H. Twelve years of SAMtools and BCFtools. Gigascience, 2021, 10(2): giab008.
[25] McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, Garimella K, Altshuler D, Gabriel S, Daly M, DePristo MA. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res, 2010, 20(9): 1297-1303.
[26] Van der Auwera GA, Carneiro MO, Hartl C, Poplin R, Del Angel G, Levy-Moonshine A, Jordan T, Shakir K, Roazen D, Thibault J, Banks E, Garimella KV, Altshuler D, Gabriel S, DePristo MA. From FastQ data to high confidence variant calls: the Genome Analysis Toolkit best practices pipeline. Curr Protoc Bioinformatics, 2013, 43(1110): 11.10. 1-11.10.33.
[27] Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MAR, Bender D, Maller J, Sklar P, de Bakker PIW, Daly MJ, Sham PC. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet, 2007, 81(3): 559-575.
[28] Treangen TJ, Salzberg SL. Repetitive DNA and next-generation sequencing: computational challenges and solutions. Nat Rev Genet, 2011, 13(1): 36-46.
[29] Chen YC, Liu T, Yu CH, Chiang TY, Hwang CC. Effects of GC bias in next-generation-sequencing data on de novo genome assembly. PLoS One, 2013, 8(4): e62856.
[30] Delahaye C, Nicolas J. Sequencing DNA with nanopores: troubles and biases. PLoS One, 2021, 16(10): e0257521.
[31] Huang Z, Xu ZX, Bai H, Huang YJ, Kang N, Ding XT, Liu J, Luo HR, Yang CT, Chen WJ, Guo QX, Xue LZ, Zhang XP, Xu L, Chen ML, Fu HG, Chen YL, Yue ZC, Fukagawa T, Liu SL, Chang GB, Xu LH. Evolutionary analysis of a complete chicken genome. Proc Natl Acad Sci USA, 2023, 120(8): e2216641120.
[32] Carson AR, Smith EN, Matsui H, Brækkan SK, Jepsen K, Hansen JB, Frazer KA. Effective filtering strategies to improve data quality from population-based whole exome sequencing studies. BMC Bioinformatics, 2014, 15: 125.
[33] Ott A, Liu SZ, Schnable JC, Yeh CTE, Wang KS, Schnable PS. tGBS® genotyping-by-sequencing enables reliable genotyping of heterozygous loci. Nucleic Acids Res, 2017, 45(21): e178.
Outlines

/