[an error occurred while processing this directive]

Hereditas(Beijing) ›› 2026, Vol. 48 ›› Issue (7): 726-740.doi: 10.16288/j.yczz.25-261

• Technique and Method • Previous Articles     Next Articles

Machine learning-based geographical ancestry inference model for the Han Chinese population

Shuaiqi Wang1,2(), Chunnian Wang1,2, Deqin Zhang2,3, Linlin Lou2,4, Yiting Ban2, Li Jiang2(), Caixia Li1,2()   

  1. 1 School of Investigation, People’s Public Security University of China, Beijing 100038, China
    2 Institute of Forensic Science, Ministry of Public Security & Beijing Engineering Research Center of Crime Scene Evidence Examination & National Engineering Laboratory for Forensic Science, Beijing 100038, China
    3 Department of Forensic Medicine, Guizhou Medical University, Guiyang 550004, China
    4 School of Forensic Medicine, Shanxi Medical University, Jinzhong 030600, China
  • Received:2025-12-23 Revised:2026-02-18 Online:2026-07-20 Published:2026-03-06
  • Contact: Li Jiang, Caixia Li E-mail:qwert0019155@gmail.com;jl@mail.bnu.edu.cn;licaixia@tsinghua.org.cn
  • Supported by:
    National Key R&D Program of China(2022YFC3341004);National Natural Science Foundation of China(82171870);Science and Technology Program of the Ministry of Public Security(2023JSZ02);Beijing Nova Program of Science and Technology(20220484149)

Abstract:

The Han Chinese population exhibits a complex genetic structure characterized by subtle yet discernible regional differentiation. Elucidating this fine-scale population structure and developing robust models for biogeographical ancestry inference are of great significance for revealing population evolutionary patterns and achieving precise ancestry inference. However, ancestry inference models specifically tailored to the genetic diversity within the domestic Han Chinese population remain scarce. In this study, we analyzed high-density SNP data from 1,229 Han Chinese individuals across eight provinces to investigate the correlation between genetic variation and geographic distribution, and to construct a machine learning-based model for regional ancestry prediction. After stringent quality control (including linkage disequilibrium pruning), we retained 208,193 SNPs for downstream analysis. Principal component analysis (PCA) and ADMIXTURE clustering revealed measurable genetic stratification corresponding to geography, supporting the delineation of seven distinct genetic clusters within the Han population. Leveraging the top principal components as features, we trained and compared multiple classifiers—XGBoost, random forest, and K-nearest neighbors—via five-fold cross-validation on the reference set, with model performance evaluated using both top-rank prediction accuracy and likelihood ratio (LR)-based metrics. The results showed that the PCA-XGBoost model achieved the optimal prediction performance in the reference set, with a first-rank prediction accuracy of 87.66% and an LR-based accuracy of 96.87%. In independent test sets, the PCA-XGBoost model maintained strong performance (first-rank prediction accuracy >85%; LR-based accuracy >95%), demonstrating excellent generalizability and stability. In summary, the PCA-XGBoost predictive model developed in this study demonstrates high efficiency, robustness, and accuracy, offering a reliable methodological tool for research in population genetics and forensic genetics.

Key words: Han Chinese, biogeographic ancestry inference, high-density SNP, principal component analysis, machine learning