Motif-aware Attribute Masking for Molecular Graph Pre-training

About

Attribute reconstruction is used to predict node or edge features in the pre-training of graph neural networks. Given a large number of molecules, they learn to capture structural knowledge, which is transferable for various downstream property prediction tasks and vital in chemistry, biomedicine, and material science. Previous strategies that randomly select nodes to do attribute masking leverage the information of local neighbors However, the over-reliance of these neighbors inhibits the model's ability to learn from higher-level substructures. For example, the model would learn little from predicting three carbon atoms in a benzene ring based on the other three but could learn more from the inter-connections between the functional groups, or called chemical motifs. In this work, we propose and investigate motif-aware attribute masking strategies to capture inter-motif structures by leveraging the information of atoms in neighboring motifs. Once each graph is decomposed into disjoint motifs, the features for every node within a sample motif are masked. The graph decoder then predicts the masked features of each node within the motif for reconstruction. We evaluate our approach on eight molecular property prediction datasets and demonstrate its advantages.

Eric Inae, Gang Liu, Meng Jiang• 2023

Related benchmarks

Task	Dataset	Result
Graph Classification	NCI1	Accuracy78.59	658
Graph Classification	NCI109	Accuracy76.82	267
Graph Classification	HIV	ROC-AUC0.7811	155
Graph property prediction	BACE	ROC AUC81.32	111
Graph property prediction	Tox21	ROC-AUC0.7829	109
Graph property prediction	ClinTox	ROC-AUC77.11	102
Graph Regression	Peptides struct (test)	MAE0.365	97
Graph property prediction	ToxCast	ROC-AUC0.6801	95
Graph property prediction	SIDER	ROC AUC62.69	95
Graph property prediction	MUV	ROC-AUC0.7241	95

Showing 10 of 30 rows

Other info

Follow for update

@wizwand_team Discord