技术与方法

基于染色体编码的多头自注意力模型进行大白猪生长性状表型的基因组预测

展开
  • 1.四川农业大学动物科技学院猪禽种业全国重点实验室成都 611130
    2.四川农业大学动物科技学院农业农村部畜禽生物组学重点实验室成都 611130
    3.四川农业大学禽遗传资源发掘与创新利用四川省重点实验室成都 611130
周文譞,硕士研究生,专业方向:遗传育种。E-mail: 1073382497@qq.com
唐国庆,教授,博士生导师,研究方向:猪遗传育种。E-mail: tyq003@163.com

收稿日期: 2025-08-20

  修回日期: 2025-10-29

  网络出版日期: 2026-01-12

基金资助

国家生猪技术创新中心先导科技项目(NCTIP-XD/B01);四川省科技厅项目(2020YFN0024);四川省科技厅项目(2021ZDZX0008);四川省科技厅项目(2021YFYZ0030);四川省猪创新团队项目(sccxtd-2022-08)

Genomic prediction of growth trait phenotypes in Large White pigs using a chromosome-encoded multi-head self-attention model

Expand
  • 1. State Key Laboratory of Swine and Poultry Breeding Industry, College of Animal Science and Technology, Sichuan Agricultural University, Chengdu 611130, China
    2. Key Laboratory of Livestock and Poultry Multi-omics of Ministry of Agriculture and Rural Affairs, College of Animal Science and Technology, Sichuan Agricultural University, Chengdu 611130, China
    3. Farm Animal Genetic Resources Exploration and Innovation Key Laboratory of Sichuan Province, Sichuan Agricultural University, Chengdu 611130, China

Received date: 2025-08-20

  Revised date: 2025-10-29

  Online published: 2026-01-12

Supported by

National Pig Technology Innovation Center Pioneering Science and Technology Project(NCTIP-XD/B01);Sichuan Provincial Department of Science and Technology Project(2020YFN0024);Sichuan Provincial Department of Science and Technology Project(2021ZDZX0008);Sichuan Provincial Department of Science and Technology Project(2021YFYZ0030);Sichuan Pig Innovation Team(sccxtd-2022-08)

摘要

随着基因组测序技术的普及,利用基因组标记预测复杂性状已成为育种关键。然而,基因组数据高维稀疏及其内部遗传标记间复杂的非线性交互特性,极大提高了精准数据分析的难度与硬件部署成本。因此本研究提出了一种基于染色体编码的多头自注意力模型(multi-head self-attention model)——ChrFormer进行基因组预测。该模型采用染色体编码器将全基因组SNP数据压缩为20个染色体特征向量和1个全局特征向量,利用多头自注意力机制动态捕获跨染色体的长程互作效应,最终通过多层感知机(multilayer perceptron,MLP)实现从基因组特征到表型的精准预测。本研究选取4,875头大白猪50K SNP基因分型数据以及4项重要生产性状(100 kg和115 kg背膘厚、100 kg和115 kg日龄)作为研究对象,采用十折交叉验证方法,以皮尔逊相关系数作为评价指标,系统比较了ChrFormer与基因组最佳线性无偏预测(genomic best linear unbiased prediction,GBLUP)、贝叶斯方法A (BayesA)和典型深度学习方法——视觉几何组(visual geometry group,VGG)、前馈神经网络(feedforward neural network,FNN)的预测性能;并且从模型参数量、训练耗时和过拟合程度等方面分析各深度学习模型的优劣。结果显示,ChrFormer在所有测试性状上的预测精度均显著优于VGG和FNN深度学习模型。在100 kg背膘、115 kg背膘和115 kg日龄这3个性状上,其预测准确度超越了传统的GBLUP和BayesA方法。虽然ChrFormer的单次迭代训练时间较长(54.88 s),但模型参数量仅约为VGG和FNN的1/10,且表现出更稳定的抗过拟合特性。本研究验证了自注意力机制的ChrFormer模型在猪生长性状表型的基因组预测的实用性,其轻量化的架构特点和稳定的预测性能,为计算资源有限的育种场开展表型的精准预测提供了切实可行的技术方法。

本文引用格式

周文譞, 赵真坚, 陈栋, 崔晟頔, 王俊戈, 陈子旸, 禹世欣, 陈佳苗, 周垚茜, 黄润杰, 唐国庆 . 基于染色体编码的多头自注意力模型进行大白猪生长性状表型的基因组预测[J]. 遗传, 2026 , 48(3) : 331 -340 . DOI: 10.16288/j.yczz.25-228

Abstract

With the widespread adoption of genome sequencing technologies, predicting complex traits using genomic markers has become a key component in breeding programs. However, the high dimensionality and sparsity of genomic data, along with the complex nonlinear interactions among genetic markers, significantly increase the difficulty of accurate data analysis and the cost of hardware deployment. Therefore, this study proposes a chromosome-encoded multi-head self-attention model, named ChrFormer, for genomic prediction. The model employs a chromosome encoder to compress whole-genome SNP data into 20 chromosome-specific feature vectors and one global feature vector. It leverages the multi-head self-attention mechanism to dynamically capture long-range interactive effects across chromosomes, and a multilayer perceptron (MLP) precisely predicts phenotype from the refined genomic features. The study selected genotyping data from 50,000 SNPs of 4,875 Large White pigs, along with four key production traits, including backfat thickness at 100 kg and 115 kg, and age at 100 kg and 115 kg. A ten-fold cross-validation approach and the Pearson correlation coefficient were used to evaluate prediction accuracy. The predictive performance of ChrFormer was systematically compared with genomic best linear unbiased prediction (GBLUP), Bayesian method A (BayesA), and representative deep learning methods, including the visual geometry group (VGG) network and the feedforward neural network (FNN). Furthermore, the study analyzed the strengths and weaknesses of each deep learning model from multiple aspects, including the number of model parameters, training time, and the extent of overfitting. The results show that ChrFormer significantly outperforms the VGG and FNN deep learning models in predictive accuracy across all tested traits. For three of the traits (backfat thickness at 100 kg and 115 kg, and days to 115 kg), its prediction accuracy surpasses that of the traditional GBLUP and BayesA methods. Although ChrFormer requires a longer training time per iteration (54.88 s), its number of parameters is only about one-tenth of that of VGG and FNN, and it demonstrates more stable resistance to overfitting. These results demonstrate that the self-attention-based ChrFormer is a practical tool for genomic phenotype prediction in animal breeding, and its lightweight architecture and stable performance offer a readily deployable solution for breeding stations with limited computational resources.

参考文献

[1] Daetwyler HD, Calus MPL, Pong-Wong R, de Los Campos G, Hickey JM. Genomic prediction in animals and plants: simulation of data, validation, reporting, and benchmarking. Genetics, 2013, 193(2): 347-365.
[2] VanRaden PM. Efficient methods to compute genomic predictions. J Dairy Sci, 2008, 91(11): 4414-4423.
[3] Yang WZ, Tempelman RJ. A Bayesian antedependence model for whole genome prediction. Genetics, 2012, 190(4): 1491-1501.
[4] Usai MG, Goddard ME, Hayes BJ. LASSO with cross- validation for genomic selection. Genet Res (Camb), 2009, 91(6): 427-436.
[5] Zhou BH, Mei BJ, Lv Q, Wang ZY, Su R. Research progress of machine learning and its application in animal genetics and breeding. China Anim Husb Vet Med, 2024, 51(12): 5348-5358.
  周铂涵, 梅步俊, 吕琦, 王志英, 苏蕊. 机器学习及其在动物遗传育种中的应用研究进展. 中国畜牧兽医, 2024, 51(12): 5348-5358.
[6] Bellot P, de Los Campos G, Pérez-Enciso M. Can deep learning improve genomic prediction of complex human traits? Genetics, 2018, 210(3): 809-819.
[7] Hayes BJ, Visscher PM, Goddard ME. Increased accuracy of artificial selection by using the realized relationship matrix. Genet Res (Camb), 2009, 91(1): 47-60.
[8] González-Camacho JM, Crossa J, Pérez-Rodríguez P, Ornella L, Gianola D. Genome-enabled prediction using probabilistic neural network classifiers. BMC Genomics, 2016, 17: 208.
[9] Lee HJ, Lee JH, Gondro C, Koh YJ, Lee SH. deepGBLUP: joint deep learning networks and GBLUP framework for accurate genomic prediction of complex traits in Korean native cattle. Genet Sel Evol, 2023, 55(1): 56.
[10] Huang ZM, Wang J, Yan ZM, Guo MZ. Differentially expressed genes prediction by multiple self-attention on epigenetics data. Brief Bioinform, 2022, 23(3): bbac117.
[11] Azodi CB, Bolger E, McCarren A, Roantree M, de Los Campos G, Shiu SH. Benchmarking parametric and machine learning models for genomic prediction of complex traits. G3 (Bethesda), 2019, 9(11): 3691-3702.
[12] Pérez P, de los Campos G. BGLR: a statistical package for whole genome regression and prediction. Genetics, 2014, 198(2): 483-495.
[13] Daetwyler HD, Pong-Wong R, Villanueva B, Woolliams JA. The impact of genetic architecture on genome-wide evaluation methods. Genetics, 2010, 185(3): 1021-1031.
[14] He KM, Zhang XY, Ren SQ, Sun J. Deep residual learning for image recognition. In: Conference on Computer Vision and Pattern Recognition (CVPR), 2016, 770-778.
[15] Gao PF, Zhao HN, Luo Z, Lin YF, Feng WJ, Li YL, Kong FJ, Li X, Fang C, Wang XT. SoyDNGP: a web-accessible deep learning framework for genomic prediction in soybean breeding. Brief Bioinform, 2023, 24(6): bbad349.
[16] Choromanski K, Likhosherstov V, Dohan D, Song X, Gane A, Sarlos T, Hawkins P, Davis J, Mohiuddin A, Kaiser L, Belanger D, Colwell L, Weller A. Rethinking Attention with Performers. In: International Conference on Learning Representations (ICLR), 2021.
[17] Rodgers JL, Nicewander WA. Thirteen ways to look at the correlation coefficient. Am Stat, 1988, 42(1): 59-66.
[18] Gianola D, Fernando RL, Stella A. Genomic-assisted prediction of genetic value with semiparametric procedures. Genetics, 2006, 173(3): 1761-1776.
[19] Mackay TFC. Epistasis and quantitative traits: using model organisms to study gene-gene interactions. Nat Rev Genet, 2014, 15(1): 22-33.
[20] Meuwissen TH, Hayes BJ, Goddard ME. Prediction of total genetic value using genome-wide dense marker maps. Genetics, 2001, 157(4): 1819-1829.
[21] Emani PS, Geradi MN, Gürsoy G, Grasty MR, Miranker A, Gerstein MB. Assessing and mitigating privacy risks of sparse, noisy genotypes by local alignment to haplotype databases. Genome Res, 2023, 33(12): 2156-2173.
[22] Cortes C, Mansour Y, Mohri M. Learning bounds for importance weighting. Adv Neural Inf Process Syst, 2010, 23: 442-450.
[23] Srivastava N, Hinton G, Krizhevsky A, Sutskever I, Salakhutdinov R. Dropout: a simple way to prevent neural networks from overfitting. J Mach Learn Res, 2014, 15(1): 1929-1958.
[24] Kitaev N, Kaiser Ł, Levskaya A. Reformer:the efficient transformer. In: International Conference on Learning Representations (ICLR), 2020.
[25] Feng XD, Yun LJ, Gao HF, Meng FJ. A review of research on attention mechanisms in machine vision. J Yunnan Minzu Univ(Nat Sci Ed), 2025, 34(4): 453-463.
  冯小丹, 云利军, 高海峰, 孟凤菊. 机器视觉注意力机制研究综述. 云南民族大学学报(自然科学版), 2025, 34(4): 453-463.
[26] Luo WJ, Li YJ, Urtasun R, Zemel R. Understanding the effective receptive field in deep convolutional neural networks. In: Annual Conference on Neural Information Processing Systems, 2016, 4356-5100.
文章导航

/