SNP density impact on kinship inference and IBS-machine learning optimization
Received date: 2025-12-08
Revised date: 2026-02-05
Online published: 2026-02-12
Supported by
Fundamental Research Funds for Institute of Forensic Science(2024JB043);Fundamental Research Funds for Institute of Forensic Science(2024JB026);National Natural Science Foundation of China(82171870);Beijing Nova Program of Science and Technology(20220484149)
In recent years, multiple panels containing varying numbers of single nucleotide polymorphisms (SNPs) have been reported in forensic genetics for kinship inference. However, systematic exploration of the impact of SNP number on inference performance and the application of machine learning algorithms remains lacking. Therefore, we evaluated the impact of SNP number on kinship inference performance and the optimization effects of machine learning methods on the identity-by-state (IBS) algorithm. We constructed multiple SNP panels with SNP numbers ranging from 15,476 to 20,838, and evaluated the performance of the likelihood ratio (LR) method and the IBS algorithm for kinship inference under different SNP numbers based on simulated pedigrees. After selecting the optimal SNP panel, we validated it using real pedigrees and further combined the IBS algorithm with machine learning methods to enhance inference performance. Our results showed that for the LR method, the sensitivity in inferring sixth and seventh degree kinships exhibited a significant positive correlation with SNP number. For the IBS algorithm, although the sensitivity in inferring fourth to seventh degree kinships showed a significant positive correlation with SNP number, the actual improvement was limited (only 0.5%~2.2% increase). Based on these results, we determined the optimal panel containing 20,838 SNPs (21K panel). The 21K panel based on the LR method could accurately infer kinships within sixth degree (with a sensitivity of 93.65% for sixth degree kinship inference), and the 21K panel based on the IBS algorithm could accurately infer kinships within third degree (with a sensitivity of 86.79% for third degree kinship inference). After combining the IBS algorithm with machine learning, the sensitivity for fourth degree kinship inference improved from 69.10% to 87.66%, the sensitivities for fifth and sixth degree kinships improved from 38.03% and 21.41% to 48.75% and 37.80%, respectively.
Key words: kinship inference; likelihood ratio method; IBS method; machine learning
Lv Dai, Zichen Tang, Zhen Jia, Li Jiang, Chuantong Zhao, Zhiyuan Zhao, Wenting Zhao, Caixia Li . SNP density impact on kinship inference and IBS-machine learning optimization[J]. Hereditas(Beijing), 2026 , 48(6) : 570 -588 . DOI: 10.16288/j.yczz.25-287
| [1] | Phillips C. The golden state killer investigation and the nascent field of forensic genealogy. Forensic Sci Int Genet, 2018, 36: 186-188. |
| [2] | Tillmar A, Kling D. Comparative study of statistical approaches and SNP panels to infer distant relationships in forensic genetics. Genes, 2025, 16(2): 114. |
| [3] | Jäger AC, Alvarez ML, Davis CP, Guzmán E, Han Y, Way L, Walichiewicz P, Silva D, Pham N, Caves G, Bruand J, Schlesinger F, Pond SJK, Varlaro J, Stephens KM, Holt CL. Developmental validation of the MiSeq FGx forensic genomics system for targeted next generation sequencing in forensic DNA casework and database laboratories. Forensic Sci Int Genet, 2017, 28: 52-70. |
| [4] | Gorden EM, Greytak EM, Sturk-Andreaggi K, Cady J, McMahon TP, Armentrout S, Marshall C. Extended kinship analysis of historical remains using SNP capture. Forensic Sci Int Genet, 2022, 57: 102636. |
| [5] | De Vries JH, Kling D, Vidaki A, Arp P, Kalamara V, Verbiest MMPJ, Piniewska-Róg D, Parsons TJ, Uitterlinden AG, Kayser M. Impact of SNP microarray analysis of compromised DNA on kinship classification success in the context of investigative genetic genealogy. Forensic Sci Int Genet, 2022, 56: 102625. |
| [6] | Tillmar A, Sturk-Andreaggi K, Daniels-Higginbotham J, Thomas JT, Marshall C. The FORCE panel: an all-in-one SNP marker set for confirming investigative genetic genealogy leads and for general forensic applications. Genes (Basel), 2021, 12(12): 1968. |
| [7] | Kling D, Tillmar A. Forensic genealogy—a comparison of methods to infer distant relationships based on dense SNP data. Forensic Sci Int Genet, 2019, 42: 113-124. |
| [8] | Kling D, Phillips C, Kennett D, Tillmar A. Investigative genetic genealogy: current methods, knowledge and practice. Forensic Sci Int Genet, 2021, 52: 102474. |
| [9] | Zhou Y, Browning SR, Browning BL. A fast and simple method for detecting identity-by-descent segments in large-scale data. Am J Hum Genet, 2020, 106(4): 426-437. |
| [10] | Seidman DN, Shenoy SA, Kim M, Babu R, Woods IG, Dyer TD, Lehman DM, Curran JE, Duggirala R, Blangero J, Williams AL. Rapid, phase-free detection of long identity- by-descent segments enables effective relationship classification. Am J Hum Genet, 2020, 106(4): 453-466. |
| [11] | Dou JZ, Sun BL, Sim XL, Hughes JD, Reilly DF, Tai ES, Liu JJ, Wang CL. Estimation of kinship coefficient in structured and admixed populations using sparse sequencing data. PLoS Genet, 2017, 13(9): e1007021. |
| [12] | Morling N, Allen RW, Carracedo A, Geada H, Guidet F, Hallenberg C, Martin W, Mayr WR, Olaisen B, Pascali VL, Schneider PM, Paternity Testing Commission of the International Society of Forensic Genetics. Paternity testing commission of the international society of forensic genetics: recommendations on genetic investigations in paternity cases. Forensic Sci Int, 2002, 129(3): 148-157. |
| [13] | Abecasis GR, Wigginton JE. Handling marker-marker linkage disequilibrium: pedigree analysis with clustered markers. Am J Hum Genet, 2005, 77(5): 754-767. |
| [14] | Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MAR, Bender D, Maller J, Sklar P, De Bakker PIW, Daly MJ, Sham PC. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet, 2007, 81(3): 559-575. |
| [15] | Zeng K, Zhao WT, Fang ZX, Li J, Liu J, Zhao D, Zhu BF, Li CX. Development and validation of a capture sequencing panel containing 9000 SNPs for inferring distant relatives in East Asian populations. Forensic Sci Int Genet, 2026, 81: 103341. |
| [16] | Abecasis GR, Cherny SS, Cookson WO, Cardon LR. Merlin—rapid analysis of dense genetic maps using sparse gene flow trees. Nat Genet, 2002, 30(1): 97-101. |
| [17] | Wu RG, Chen H, Li R, Zang Y, Shen XF, Hao B, Wang QW, Sun HY. Pairwise kinship testing with microhaplotypes: can advancements be made in kinship inference with these markers? Forensic Sci Int, 2021, 325: 110875. |
| [18] | Cui W, Chen M, Yang Y, Cai MM, Lan Q, Xie T, Zhu BF. Applications of 1993 single nucleotide polymorphism loci in forensic pairwise kinship identifications and inferences. Forensic Sci Int Genet, 2023, 65: 102889. |
| [19] | Manichaikul A, Mychaleckyj JC, Rich SS, Daly K, Sale M, Chen WM.Robust relationship inference in genome-wide association studies. Bioinformatics, 2010, 26(22): 2867-2873. |
| [20] | Liu J, Wei YL, Yang L, Jiang L, Zhao WT, Li CX. Testing of two SNP array-based genealogy algorithms using extended Han Chinese pedigrees and recommendations for improved performances in forensic practice. Electrophoresis, 2023, 44(17-18): 1435-1445. |
| [21] | Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: synthetic minority over-sampling technique. J Artif Intell Res, 2002, 16(1): 321-357. |
| [22] | Van Den Heuvel E, Zhan ZZ. Myths about linear and monotonic associations: Pearson’s r, Spearman’s ρ, and Kendall’s τ. Am Stat, 2022, 76(1): 44-52. |
| [23] | Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc B, 1995, 57(1): 289-300. |
| [24] | Tomczak M, Tomczak E. The need to report effect size estimates revisited. An overview of some recommended measures of effect size. Trends Sport Sci, 2014, 1(21): 19-25. |
| [25] | Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Front Psychol, 2013, 4: 863. |
| [26] | Hill WG, Weir BS. Variation in actual relationship as a consequence of Mendelian sampling and linkage. Genet Res, 2011, 93(1): 47-64. |
| [27] | Huff CD, Witherspoon DJ, Simonson TS, Xing JC, Watkins WS, Zhang YH, Tuohy TM, Neklason DW, Burt RW, Guthery SL, Woodward SR, Jorde LB. Maximum- likelihood estimation of recent shared ancestry (ERSA). Genome Res, 2011, 21(5): 768-774. |
| [28] | Speed D, Balding DJ. Relatedness in the post-genomic era: is it still useful? Nat Rev Genet, 2015, 16(1): 33-44. |
| [29] | Wei YF, Zhu Q, Wang HY, Cao YY, Li X, Zhang XK, Wang YF, Zhang J. Pairwise kinship inference and pedigree reconstruction using 91 microhaplotypes. Forensic Sci Int Genet, 2024, 72: 103090. |
| [30] | Tabangin ME, Woo JG, Martin LJ. The effect of minor allele frequency on the likelihood of obtaining false positives. BMC Proc, 2009, 3(Suppl 7): S41. |
| [31] | Thompson EA. Identity by descent: variation in meiosis, across genomes, and in populations. Genetics, 2013, 194(2): 301-326. |
/
| 〈 |
|
〉 |