遗传 ›› 2026, Vol. 48 ›› Issue (9): 931-945.doi: 10.16288/j.yczz.25-275
张巧生1(
), 许俊杰1, 孙振宇1, 仲兆满1, 刘婕1, 吴延立2, 李婉琴1, 胡梦杰1, 李鸿鹏3(
)
收稿日期:2025-12-17
修回日期:2026-02-21
出版日期:2026-04-17
发布日期:2026-04-17
通讯作者:
李鸿鹏,博士,讲师,研究方向:机器学习、微分方程数值解。E-mail: 2019000027@jou.edu.cn作者简介:张巧生,博士,副教授,研究方向:基因组学大数据、数量性状遗传解析、生物信息学。E-mail: zqs@jou.edu.cn
基金资助:
Qiaosheng Zhang1(
), Junjie Xu1, Zhenyu Sun1, Zhaoman Zhong1, Jie Liu1, Yanli Wu2, Wanqin Li1, Mengjie Hu1, Hongpeng Li3(
)
Received:2025-12-17
Revised:2026-02-21
Published:2026-04-17
Online:2026-04-17
Supported by:摘要:
丰富的组学数据极大地推动了多组学数据整合技术的发展,采用非线性嵌入方式整合数据在多组学研究中逐渐成为主流,这种方式能够通过提升嵌入的质量以显著改善对癌症的分析。然而,现有的多组学数据整合方法通常局限于组学测量,忽略了包括生物学通路在内的领域先验知识。本文提出了一种基于通路自注意力和图卷积网络(graph convolutional network,GCN)的多组学整合分类模型——PathTransGCN,该模型将生物学通路信息整合到多组学数据分析中,旨在提高癌症分类的准确性。从癌症基因组图谱(The Cancer Genome Atlas,TCGA)和UCSC Xena数据库中获取乳腺癌(breast cancer,BRCA)、非小细胞肺癌(non-small cell lung cancer,NSCLC)和低级别脑胶质瘤(low-grade glioma,LGG)的多组学数据,包括基因突变、DNA甲基化、拷贝数变异和基因表达,以评估模型在不同癌症中的泛化能力。首先,PathTransGCN利用通路自注意力模块学习样本在不同通路中的潜在表征,得到多组学整合向量。同时,基于相似性网络融合(similarity network fusion,SNF)方法构建患者样本相似性网络(patient similarity network,PSN)。然后,将整合后的向量和PSN共同输入到GCN中进行端到端训练,实现癌症亚型的精准分类。在对BRCA数据集的多组学数据进行分析时,PathTransGCN在癌症亚型五分类中的性能高于目前流行的几种算法(如MoGCN、DeePathNet等),准确率和F1分数分别达到87.6%和86.4%。此外,该模型在NSCLC和LGG数据集上均展现出良好的泛化能力,并能在通路水平上有效识别与疾病相关的关键生物标志物。实验结果表明,PathTransGCN在组学数据整合和分类结果的可解释性方面表现出色,在临床应用方面具有巨大潜力。
张巧生, 许俊杰, 孙振宇, 仲兆满, 刘婕, 吴延立, 李婉琴, 胡梦杰, 李鸿鹏. 基于通路自注意力和图卷积网络的非线性多组学数据整合分类模型[J]. 遗传, 2026, 48(9): 931-945.
Qiaosheng Zhang, Junjie Xu, Zhenyu Sun, Zhaoman Zhong, Jie Liu, Yanli Wu, Wanqin Li, Mengjie Hu, Hongpeng Li. A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks[J]. Hereditas(Beijing), 2026, 48(9): 931-945.
表2
BRCA数据集的分类结果"
| 方法 | 准确率 | 加权F1 | 宏F1 | 精度 | 召回率 |
|---|---|---|---|---|---|
| RF | 0.768±0.032 | 0.756±0.031 | 0.697±0.029 | 0.731±0.028 | 0.675±0.033 |
| KNN | 0.783±0.027 | 0.777±0.026 | 0.732±0.025 | 0.801±0.024 | 0.692±0.031 |
| LR | 0.772±0.030 | 0.752±0.032 | 0.709±0.030 | 0.792±0.027 | 0.672±0.034 |
| XGBoost | 0.791±0.025 | 0.786±0.024 | 0.730±0.026 | 0.775±0.029 | 0.700±0.028 |
| SVM | 0.788±0.032 | 0.783±0.027 | 0.728±0.027 | 0.779±0.026 | 0.739±0.025 |
| DeePathNet | 0.855±0.018 | 0.841±0.047 | 0.811±0.039 | 0.833±0.018 | 0.807±0.027 |
| MoGCN | 0.837±0.031 | 0.834±0.030 | 0.798±0.022 | 0.842±0.029 | 0.770±0.033 |
| PathTransGCN | 0.876±0.029 | 0.864±0.021 | 0.832±0.027 | 0.858±0.018 | 0.796±0.021 |
表3
NSCLC数据集的分类结果"
| 方法 | 准确率 | 曲线下面积 | F1分数 | 精度 | 召回率 |
|---|---|---|---|---|---|
| RF | 0.813±0.031 | 0.830±0.021 | 0.812±0.023 | 0.881±0.024 | 0.802±0.033 |
| KNN | 0.695±0.022 | 0.717±0.023 | 0.709±0.022 | 0.696±0.021 | 0.763±0.024 |
| LR | 0.805±0.036 | 0.812±0.031 | 0.795±0.034 | 0.812±0.032 | 0.867±0.025 |
| XGBoost | 0.857±0.031 | 0.861±0.022 | 0.849±0.021 | 0.903±0.020 | 0.835±0.021 |
| SVM | 0.836±0.025 | 0.842±0.020 | 0.828±0.023 | 0.887±0.024 | 0.817±0.022 |
| DeePathNet | 0.885±0.043 | 0.889±0.028 | 0.884±0.020 | 0.926±0.023 | 0.832±0.020 |
| MoGCN | 0.873±0.023 | 0.877±0.031 | 0.864±0.023 | 0.915±0.022 | 0.849±0.021 |
| PathTransGCN | 0.923±0.027 | 0.928±0.021 | 0.924±0.022 | 0.953±0.014 | 0.861±0.022 |
表4
LGG数据集的分类结果"
| 方法 | 准确率 | 加权F1 | 宏F1 | 精度 | 召回率 |
|---|---|---|---|---|---|
| RF | 0.792±0.028 | 0.785±0.026 | 0.741±0.024 | 0.773±0.025 | 0.752±0.029 |
| KNN | 0.768±0.031 | 0.761±0.029 | 0.712±0.030 | 0.754±0.028 | 0.731±0.032 |
| LR | 0.785±0.029 | 0.776±0.030 | 0.735±0.027 | 0.782±0.026 | 0.740±0.031 |
| XGBoost | 0.815±0.024 | 0.809±0.022 | 0.768±0.023 | 0.803±0.025 | 0.775±0.026 |
| SVM | 0.807±0.027 | 0.801±0.025 | 0.762±0.026 | 0.798±0.024 | 0.782±0.023 |
| DeePathNet | 0.862±0.020 | 0.853±0.022 | 0.821±0.019 | 0.847±0.018 | 0.815±0.021 |
| MoGCN | 0.848±0.025 | 0.842±0.023 | 0.809±0.021 | 0.851±0.022 | 0.798±0.024 |
| PathTransGCN | 0.894±0.022 | 0.887±0.019 | 0.856±0.020 | 0.882±0.017 | 0.843±0.020 |
| [1] |
Chen CY, Wang J, Pan DH, Wang XY, Xu YP, Yan JJ, Wang LZ, Yang XF, Yang M, Liu GP. Applications of multi-omics analysis in human diseases. MedComm (2020), 2023, 4(4): e315.
pmid: 37533767 |
| [2] |
Hayes CN, Nakahara H, Ono A, Tsuge M, Oka S. From omics to multi-omics: a review of advantages and tradeoffs. Genes (Basel), 2024, 15(12): 1551.
pmid: 39766818 |
| [3] | Li ZH, Chen MQ, Yuan XJ, Huang HJ, Huang WT, Zhou PY, Zeng C, Feng XN, Yang LY, Huang SQ, Tan CY, Chen CR, Yan QX. Identification of biomarkers for non- obstructive azoospermia based on microRNA and bioinformatics screening. Hereditas(Beijing), 2026, 48(3): 301-312. |
| 李志宏, 陈淼琪, 袁晓珺, 黄华君, 黄琬婷, 周飘雁, 曾晨, 冯许诺, 杨洛瑶, 黄树强, 谭翠钰, 陈彩蓉. 基于microRNA与生物信息学筛选非梗阻性无精子症的生物标志物. 遗传, 2026, 48(3): 301-312. | |
| [4] | Ismaeel AG, Mikhail DY. Effective data mining technique for classification cancers via mutations in gene using neural network. arXiv, 2016, doi: 10.48550/arXiv.1608.02888. |
| [5] |
Jiang Q, Jin M. Feature selection for breast cancer classification by integrating somatic mutation and gene expression. Front Genet, 2021, 12: 629946.
pmid: 33719339 |
| [6] |
Sallis BF, Erkert L, Moñino-Romero S, Acar U, Wu RN, Konnikova L, Lexmond WS, Hamilton MJ, Dunn WA, Szepfalusi Z, Vanderhoof JA, Snapper SB, Turner JR, Goldsmith JD, Spencer LA, Nurko S, Fiebiger E. An algorithm for the classification of mRNA patterns in eosinophilic esophagitis: integration of machine learning. J Allergy Clin Immunol, 2018, 141(4): 1354-1364.e9.
pmid: 29273402 |
| [7] | Xie BB, Yang YD, Ding N, Fang XD. Identification of disease targets for precision medicine by integrative analysis of multi-omics data. Hereditas(Beijing), 2015, 37(7): 655-663. |
| 谢兵兵, 杨亚东, 丁楠, 方向东. 整合分析多组学数据筛选疾病靶点的精准医学策略. 遗传, 2015, 37(7): 655-663. | |
| [8] |
Huang SJ, Chaudhary K, Garmire LX. More is better: recent progress in multi-omics data integration methods. Front Genet, 2017, 8: 84.
pmid: 28670325 |
| [9] |
Fiocchi C. Omics and multi-omics in IBD: no integration, no breakthroughs. Int J Mol Sci, 2023, 24(19): 14912.
pmid: 37834360 |
| [10] | Abdi H, Williams LJ. Principal component analysis. Wires Comput Stat, 2010, 2(4): 433-459. |
| [11] | McInnes L, Healy J, Melville J. Umap: uniform manifold approximation and projection for dimension reduction. arXiv, 2018, doi: 10.48550/arXiv.1802.03426. |
| [12] | Ng A. Sparse autoencoder. CS294A Lecture Notes, 2011, 72(2011): 1-19. |
| [13] |
Makrodimitris S, Pronk B, Abdelaal T, Reinders M. An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics. Brief Bioinform, 2024, 25(1): bbad416.
pmid: 38018908 |
| [14] |
Park M, Kim D, Moon K, Park T. Integrative analysis of multi-omics data based on blockwise sparse principal components. Int J Mol Sci, 2020, 21(21): 8202.
pmid: 33147797 |
| [15] |
Picard M, Scott-Boyer MP, Bodein A, Périn O, Droit A. Integration strategies of multi-omics data for machine learning analysis. Comput Struct Biotechnol J, 2021, 19: 3735-3746.
pmid: 34285775 |
| [16] |
Chaudhary K, Poirion OB, Lu LQ, Garmire LX. Deep learning-based multi-omics integration robustly predicts survival in liver cancer. Clin Cancer Res, 2018, 24(6): 1248-1259.
pmid: 28982688 |
| [17] |
Li X, Ma J, Leng L, Han MF, Li MS, He FC, Zhu YP. MoGCN: a multi-omics integration method based on graph convolutional network for cancer subtype analysis. Front Genet, 2022, 13: 806842.
pmid: 35186034 |
| [18] |
Tan KW, Huang WX, Hu JL, Dong SB. A multi-omics supervised autoencoder for pan-cancer clinical outcome endpoints prediction. BMC Med Inform Decis Mak, 2020, 20(Suppl 3): 129.
pmid: 32646413 |
| [19] | Zhang ZY, Wang QL, Zhang JY, Duan YY, Liu JX, Liu ZS, Li CY. Machine learning applications in breast cancer survival and therapeutic outcome prediction based on multi-omic analysis. Hereditas(Beijing), 2024, 46(10): 820-832. |
| 章子怡, 王棨临, 张俊有, 段迎迎, 刘家欣, 刘赵硕, 李春燕. 多组学数据驱动的机器学习模型在乳腺癌生存及治疗响应预测中的应用. 遗传, 2024, 46(10): 820-832. | |
| [20] |
Yu TW. AIME: autoencoder-based integrative multi-omics data embedding that allows for confounder adjustments. PLoS Comput Biol, 2022, 18(1): e1009826.
pmid: 35081109 |
| [21] |
Tan CY, Ong HF, Lim CH, Tan MS, Ooi EH, Wong K. Amogel: a multi-omics classification framework using associative graph neural networks with prior knowledge for biomarker identification. BMC Bioinformatics, 2025, 26(1): 94.
pmid: 40155814 |
| [22] |
Yan HX, Weng DW, Li DG, Gu Y, Ma WJ, Liu QJ. Prior knowledge-guided multilevel graph neural network for tumor risk prediction and interpretation via multi-omics data integration. Brief Bioinform, 2024, 25(3): bbae184.
pmid: 38670157 |
| [23] |
Lan W, Liao HB, Chen QF, Zhu LZ, Pan Y, Chen YPP. DeepKEGG: a multi-omics data integration framework with biological insights for cancer recurrence prediction and biomarker discovery. Brief Bioinform, 2024, 25(3): bbae185.
pmid: 38678587 |
| [24] |
Wang B, Mezlini AM, Demir F, Fiume M, Tu ZW, Brudno M, Haibe-Kains B, Goldenberg A. Similarity network fusion for aggregating data types on a genomic scale. Nat Methods, 2014, 11(3): 333-337.
pmid: 24464287 |
| [25] | He KM, Chen XL, Xie SN, Li YH, Dollár P, Girshick R. Masked autoencoders are scalable vision learners. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, LA, USA, 2022, 15979-15988. |
| [26] |
Wei L, Jin ZL, Yang SJ, Xu YX, Zhu YT, Ji Y. TCGA- assembler 2: software pipeline for retrieval and processing of TCGA/CPTAC data. Bioinformatics, 2018, 34(9): 1615-1617.
pmid: 29272348 |
| [27] |
Orrantia-Borunda E, Anchondo-Nuñez P, Acuña-Aguilar LE, Gómez-Valles FO, Ramírez-Valdespino CA. Subtypes of breast cancer. Breast Cancer-tokyo, 2022, doi: 10.36255/exon-publications-breast-cancer-subtypes.
pmid: 36122153 |
| [28] |
Herbst RS, Morgensztern D, Boshoff C. The biology and management of non-small cell lung cancer. Nature, 2018, 553(7689): 446-454.
pmid: 29364287 |
| [29] |
Kuenzi BM, Ideker T. A census of pathway maps in cancer systems biology. Nat Rev Cancer, 2020, 20(4): 233-246.
pmid: 32066900 |
| [30] | Naidu G, Zuva T, Sibanda EM. A review of evaluation metrics in machine learning algorithms. Computer Science On-line Conference, 2023, 15-25. |
| [31] | Bradley AP. The use of the area under the ROC curve in the evaluation of machine learning algorithms. Pattern Recogn, 1997, 30(7): 1145-1159. |
| [32] | Liao LZ, Li H, Shang WY, Ma L. An empirical study of the impact of hyperparameter tuning and model optimization on the performance properties of deep neural networks. Acm T Softw Eng Meth, 2022, 31(3): 1-40. |
| [33] | Loshchilov I, Hutter F. Decoupled weight decay regularization. arXiv, 2017, doi: 10.48550/arXiv.1711.05101. |
| [34] |
Zhou P, Xie XY, Lin ZC, Yan SC. Towards understanding convergence and generalization of AdamW. IEEE Trans Pattern Anal Mach Intell, 2024, 46(9): 6486-6493.
pmid: 38536692 |
| [35] | Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I. Attention is all you need. arXiv, 2017, doi: 10.48550/arXiv.1706.03762. |
| [36] |
Bao JX, Chang C, Zhang QYW, Saykin AJ, Shen L, Long Q, Alzheimer’s Disease Neuroimaging Initiative. Integrative analysis of multi-omics and imaging data with incorporation of biological information via structural Bayesian factor analysis. Brief Bioinform, 2023, 24(2): bbad073.
pmid: 36882008 |
| [37] |
Cantini L, Zakeri P, Hernandez C, Naldi A, Thieffry D, Remy E, Baudot A. Benchmarking joint multi-omics dimensionality reduction approaches for the study of cancer. Nat Commun, 2021, 12(1): 124.
pmid: 33402734 |
| [38] |
Dennis G Jr, Sherman BT, Hosack DA, Yang J, Gao W, Lane HC, Lempicki RA. DAVID: database for annotation, visualization, and integrated discovery. Genome Biol, 2003, 4(5): P3.
pmid: 12734009 |
| [39] |
Lu B, Qiu R, Wei JT, Wang L, Zhang QK, Li MS, Zhan XD, Chen J, Hsieh IY, Yang CQ, Zhang J, Sun ZC, Zhu YF, Jiang T, Zhu H, Li J, Zhao W. Phase separation of phospho-HDAC6 drives aberrant chromatin architecture in triple-negative breast cancer. Nat Cancer, 2024, 5(11): 1622-1640.
pmid: 39198689 |
| [40] |
Hammer A, Diakonova M. Tyrosyl phosphorylated serine- threonine kinase PAK1 is a novel regulator of prolactin- dependent breast cancer cell motility and invasion. Adv Exp Med Biol, 2015, 846: 97-137.
pmid: 25472536 |
| [41] |
Pan LH, Li JL, Xu Q, Gao ZL, Yang M, Wu XP, Li XS. HER2/PI3K/AKT pathway in HER2-positive breast cancer: a review. Medicine (Baltimore), 2024, 103(24): e38508.
pmid: 38875362 |
| [42] |
Dawson MA, Bannister AJ, Göttgens B, Foster SD, Bartke T, Green AR, Kouzarides T. JAK2 phosphorylates histone H3Y41 and excludes HP1α from chromatin. Nature, 2009, 461(7265): 819-822.
pmid: 19783980 |
| [43] |
Lin NU, Winer EP. New targets for therapy in breast cancer: small molecule tyrosine kinase inhibitors. Breast Cancer Res, 2004, 6(5): 204-210.
pmid: 15318926 |
| [44] |
Adhikari VP, Lu LJ, Kong LQ. Does hepatitis B virus infection cause breast cancer? Chin Clin Oncol, 2016, 5(6): 81.
pmid: 27701872 |
| [45] |
Lee JS, Tocheny CE, Shaw LM. The insulin-like growth factor signaling pathway in breast cancer: an elusive therapeutic target. Life (Basel), 2022, 12(12): 1992.
pmid: 36556357 |
| [1] | 丰继华, 陈忠兴, 康琦林, 李龙飞, 杨佳慧, 张雨亭. 融合通道与空间注意力机制的转录因子结合位点预测方法[J]. 遗传, 2026, 48(5): 522-534. |
| [2] | 周文譞, 赵真坚, 陈栋, 崔晟頔, 王俊戈, 陈子旸, 禹世欣, 陈佳苗, 周垚茜, 黄润杰, 唐国庆. 基于染色体编码的多头自注意力模型进行大白猪生长性状表型的基因组预测[J]. 遗传, 2026, 48(3): 331-340. |
| [3] | 高炳熙, 吴华煊, 杜志强. 应用图像转换与深度学习提升单细胞分类精度[J]. 遗传, 2025, 47(3): 382-392. |
| [4] | 鲍艳春, 石彩霞, 张传强, 谷明娟, 朱琳, 刘在霞, 周乐, 马凤英, 娜日苏, 张文广. 深度学习在基因组学中的研究进展[J]. 遗传, 2024, 46(9): 701-715. |
| [5] | 杨帆, 韩巧玲, 赵文迪, 赵玥. 基于层级和全局特征结合的蛋白质序列EC编号预测[J]. 遗传, 2024, 46(8): 661-669. |
| [6] | 章子怡, 王棨临, 张俊有, 段迎迎, 刘家欣, 刘赵硕, 李春燕. 多组学数据驱动的机器学习模型在乳腺癌生存及治疗响应预测中的应用[J]. 遗传, 2024, 46(10): 820-832. |
| [7] | 郑慧怡, 吴华煊, 杜志强. 肠道宏基因组图像增强和深度学习改善代谢性疾病分类预测精度[J]. 遗传, 2024, 46(10): 886-896. |
| [8] | 胡伟澎, 李佑平, 张秀清. 基于迁移学习的MHC-I型抗原表位呈递预测[J]. 遗传, 2019, 41(11): 1041-1049. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||
www.chinagene.cn
备案号:京ICP备09063187号-4
总访问:,今日访问:,当前在线: