Item

Learning Expressive, Explainable and Generalizable Representations for Graph-structured Data

Yao, Tianjun
Citations
Altmetric:
Supervisor
Department
Machine Learning
Embargo End Date
Type
Dissertation
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Graph Neural Networks (GNNs) have emerged as a powerful paradigm for learning representations from graph-structured data, achieving remarkable success across diverse domains including molecular property prediction, social network analysis, and recommendation systems. The central challenge in graph representation learning lies in how to learn high-quality representations for graph-structured data. We define high-quality graph representations through three essential properties: (1) expressiveness: the ability to distinguish structurally different graphs and capture fine-grained substructure information; (2) explainability: the capacity to identify interpretable substructures that causally drive predictions; and (3) generalizability: the robustness to maintain predictive performance under distribution shifts between training and test data. This dissertation addresses these challenges by developing novel frameworks. For expressiveness, we first introduce SEK-GNN, a random walk-based structural encoding framework that provably enhances GNN expressiveness beyond the 1-WL test by capturing substructure counting capabilities. We then propose MuGSI, a multi-granularity knowledge distillation framework that transfers expressive structural information from teacher GNNs to efficient student MLPs at graph, subgraph, and node levels, achieving improved expressivity while maintaining computational efficiency. For explainability and generalizability, we develop a causality-inspired framework grounded in invariant learning principles. We introduce EQuAD, which leverages the infomax principle to first learn spurious features and then disentangle them from invariant features, enabling both interpretable subgraph identification and robust out-of-distribution (OOD) generalization. Building upon EQuAD, we propose LIRS, which employs biased infomax to more effectively capture spurious patterns and a class-conditioned cross-entropy loss to enhance invariant feature learning. Finally, we present PrunE, a pruning-based approach that achieves OOD generalization by removing spurious edges rather than directly identifying invariant subgraphs, representing a unification of explainability, generalizability, and expressivity in a single unified framework.
Citation
Yao, Tianjun, "Learning Expressive, Explainable and Generalizable Representations for Graph-structured Data," PhD Dissertation, Machine Learning, MBZUAI, 2026.
Source
Conference
Keywords
out-of-distribution generalization, Graph neural networks, expressive power, explainability, deep learning, machine learning
Subjects
Source
Publisher
DOI
Full-text link