遗传

• 技术与方法 •    

基于通路自注意图卷积网络的非线性多组学数据整合分类模型

张巧生1,许俊杰1,孙振宇1,仲兆满1,刘婕1,吴延立2,李婉琴1,胡梦杰1,李鸿鹏3   

  1. 1.江苏海洋大学计算机工程学院,连云港 222005

    2.山东省菏泽市食品药品检验检测研究院,菏泽 274000

    3.江苏海洋大学理学院,连云港 222005

  • 收稿日期:2025-12-17 修回日期:2026-02-21 出版日期:2026-04-17 发布日期:2026-04-17
  • 通讯作者: 李鸿鹏,博士,讲师,研究方向:机器学习、微分方程数值解。E-mail: 2019000027@jou.edu.cn
  • 基金资助:

    国家自然科学基金(编号:72174079),连云港市科技计划项目(编号:CG2323)和连云港市博士后科研资助计划(编号:LYG20210010) [ Supported by the National Natural Science Foundation of China (No. 72174079), the Lianyungang Science and Technology Planning Project (No. CG2323), and the Lianyungang Postdoctoral Research Funding Program (No. LYG20210010)]

A nonlinear multi-omics data integration and classification model based on pathway self-Attention and graph convolutional networks

Qiaosheng Zhang1,Junjie Xu1,Zhenyu Sun1,Zhaoman Zhong1,Jie Liu1,Yanli Wu2,Wanqin Li1,Mengjie Hu1,Hongpeng Li3   

  1. 1. School of Computer Engineering, Jiangsu Ocean University, Lianyungang 222005, China
    2. Shandong Heze Institute for Food and Drug Control, Heze 274000, China 
    3. School of Science, Jiangsu Ocean University, Lianyungang 222005, China
  • Received:2025-12-17 Revised:2026-02-21 Published:2026-04-17 Online:2026-04-17

摘要:

丰富的组学数据极大推动了多组学数据整合技术的发展,采用非线性嵌入方式整合数据在多组学研究中逐渐成为主流,这种方式能够通过提升嵌入的质量以显著改善癌症的分析。然而,现有的多组学数据整合方法通常局限于组学测量,忽略了包括生物学通路在内的领域先验知识。本文提出了一种基于通路自注意图卷积网络graph convolutional network, GCN的多组学整合分类模型——PathTransGCN,该模型将生物学通路信息整合到多组学数据分析中,旨在提高癌症分类的准确性。从癌症基因组图谱(The Cancer Genome Atlas, TCGA和UCSC Xena数据库中获取乳腺癌(breast cancer, BRCA)、非小细胞肺癌(non-small cell lung cancer, NSCLC)和低级别脑胶质瘤(low-grade glioma, LGG)的多组学数据,包括基因突变、DNA甲基化、拷贝数变异和基因表达,以评估模型在不同癌症下的泛化能力。首先PathTransGCN利用通路自注意力模块学习样本在不同通路中的潜在表征,得到多组学整合向量。同时,基于相似性网络融合(similarity network fusion, SNF)方法构建患者样本相似性网络(patient similarity network, PSN)。然后,整合后的向量和PSN共同输入到GCN中进行端到端训练,实现癌症亚型的精准分类。在对BRCA数据集的多组学数据进行分析时,PathTransGCN在癌症亚型分类中的性能高于目前流行的几种算法(如MoGCNDeePathNet等),准确率和F1分数分别达到87.6%86.4%。此外,该模型在NSCLCLGG数据集上均展现出良好的泛化能力,并能在通路水平上有效识别与疾病相关的关键生物标志物。实验结果表明,PathTransGCN组学数据整合和分类结果的可解释性方面表现出色,在临床应用方面具有巨大潜力。

关键词: 多组学数据整合, 生物学通路, 相似性网络融合, 深度学习, 癌症亚型分类

Abstract:

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, dna methylation, copy number variations, and gene expression, and were used to assess the model’s generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Key words:

multi-omics integration; biological pathway, similarity network fusion, deep learning, cancer subtype classification