[an error occurred while processing this directive]

Hereditas(Beijing) ›› 2026, Vol. 48 ›› Issue (9): 931-945.doi: 10.16288/j.yczz.25-275

• Technique and Method • Previous Articles     Next Articles

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks

Qiaosheng Zhang1(), Junjie Xu1, Zhenyu Sun1, Zhaoman Zhong1, Jie Liu1, Yanli Wu2, Wanqin Li1, Mengjie Hu1, Hongpeng Li3()   

  1. 1 School of Computer Engineering, Jiangsu Ocean University, Lianyungang 222005, China
    2 Heze Institute for Food and Drug Control, Shandong Province, Heze 274000, China
    3 School of Science, Jiangsu Ocean University, Lianyungang 222005, China
  • Received:2025-12-17 Revised:2026-02-21 Online:2026-04-17 Published:2026-04-17
  • Contact: Hongpeng Li E-mail:zqs@jou.edu.cn;2019000027@jou.edu.cn
  • Supported by:
    National Natural Science Foundation of China(72174079);Lianyungang Science and Technology Planning Project(CG2323);Lianyungang Postdoctoral Research Funding Program(LYG20210010)

Abstract:

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model’s generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Key words: multi-omics integration, biological pathway, similarity network fusion, deep learning, cancer subtype classification