遗传

• 研究报告 •    

基于1240K SNPs组合的法医SNP系谱推断研究

汤子琛1,2,贾镇3,方之晓2,姬鹏超1,2,江丽2,赵雯婷2,刘京2,魏以梁1,李彩霞2   

  1. 1江苏师范大学生命科学学院,苏省系统发育与比较基因组学重点实验室,徐州 221116

    2 公安部物证鉴定中心,法医遗传学公安部重点实验室,现场物证溯源技术国家工程实验室,北京 100038

    3中国人民公安大学侦查学院,北京 100038

  • 发布日期:2026-09-07
  • 基金资助:
    公安部鉴定中心基本科研业务费专项资金项目(编号:2024JB026,2024JB043),国家自然科学基金项目(编号:81772027),国家重点研发计划(编号:81772027),北京市科技新星计划(编号:20220484149)资助[Supported by the Basic Research Projects of Institute of Forensic Science (Nos. 2024JB026, 2024JB043), the National Natural Science Foundation of China (No. 81772027), the National Key R&D Program of China (No. 2022YFC3341004), and Beijing Nova Program (No. 20220484149)]

Forensic investigative genetic genealogy research based on the 1240K SNPs panel

Zichen Tang1,2Zhen Jia3Zhixiao Fang2, Pengchao Ji1,2Wenting Zhao2Jing Liu2Yiliang Wei1Caixia Li2   

  1. 1Jiangsu Key Laboratory of Phylogenomics and Comparative Genomics, School of Life Sciences, Jiangsu Normal University, Xuzhou 221000, China

    2 Key Laboratory of Forensic Genetics, Beijing Engineering Research Center of Crime Scene Evidence Examination, National Engineering Laboratory for Forensic Science, Institute of Forensic Science, Beijing 100038, China

    3 School of Criminal Investigation, People’s Public Security University of China, Beijing 100038, China

  • Online:2026-09-07

摘要: 法医SNP系谱推断技术通过分析高密度SNP中的共祖片段(identity by descent,IBD)推断亲缘关系,已成为疑难案件侦破的关键手段,但现有SNP芯片对DNA质量要求高,且IBD算法参数多基于欧美人群,在东亚人群中缺乏系统验证。针对上述问题,本研究引入古DNA领域的1240K SNP组合,建立了基于IBD算法的远距离亲缘关系推断分析流程。利用千人基因组计划504例东亚人群样本评估1240K位点组合的遗传多态性,基于161例中国汉族志愿者真实家系(1,3211~9级亲缘关系)评估IBD算法效能,并系统优化了IBD长度阈值、错配参数及连锁不平衡相关系数(R²)。结果表明:1240K SNPs在东亚人群中表现出良好的遗传多样性与群体区分能力;当IBD长度阈值优化为2 cM、错配参数优化为het13hom12时,亲缘推断能力显著提升(7级置信区间准确率CIA83.70%Macro-F1较默认参数提高2.9%),而R²过滤对算法优化作用不显著。优化后算法推断7级和8级亲缘关系的CIA分别为83.70%74.19%。本研究证实,1240K SNPs不仅可用于人群遗传结构分析,还可实现8级以内亲缘关系的可靠推断,在法医遗传学领域具有明确的应用前景,为微量、降解DNA的系谱分析提供了可行方案。

关键词: 1240K , SNPs 组合, IBD算法, 东亚人群, 系谱推断

Abstract: Forensic SNP genealogy inference identifies kinship by analyzing identity-by-descent (IBD) segments in high-density SNP data and has become a key technique for solving cold cases. However, current methods face two major challenges: (1) SNP microarrays require high-quality DNA, making it difficult to analyze trace or degraded DNA commonly found at crime scenes; (2) existing IBD algorithm parameters are mainly calibrated on European populations, lacking systematic validation in East Asian populations. To address these issues, we introduced the 1240K single nucleotide polymorphism (SNP) panel originally developed for ancient DNA research and established a long-range kinship inference pipeline based on the IBD algorithm. We evaluated the genetic polymorphism of the 1240K SNP panel using 504 East Asian individuals from the 1000 Genomes Project, assessed the performance of the IBD algorithm on 1,321 pairwise relationships (1st to 9th degree) derived from 161 real Chinese Han pedigrees, and systematically optimized three key parameters: IBD length threshold (cM), mismatch tolerance (het/hom), and linkage disequilibrium coefficient (R²). The results showed that the 1240K SNPs exhibited favorable genetic diversity and population discrimination power in East Asians. Optimizing the IBD length threshold to 2 cM and mismatch parameters to het13hom12 significantly improved kinship inference accuracy (7th degree confidence interval accuracy reached 83.70%, and Macro-F1 increased by 2.9% compared with default settings), whereas R²-based filtering did not substantially enhance performance. Under the optimized parameters, the 7th and 8th degree confidence interval accuracies were 83.70% and 74.19%, respectively. This study confirms that the 1240K SNP panel can be used not only for population genetic structure analysis but also for reliable kinship inference up to the 8th degree, demonstrating clear application potential in forensic genetics. It provides a feasible solution for genealogical analysis of trace and degraded DNA.

Key words:  , IBD algorithm,  , 1240K SNPs,  , East Asian , population,  , kinship inference