Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Do Transformers Really Perform Bad for Graph Representation?

About

The Transformer architecture has become a dominant choice in many domains, such as natural language processing and computer vision. Yet, it has not achieved competitive performance on popular leaderboards of graph-level prediction compared to mainstream GNN variants. Therefore, it remains a mystery how Transformers could perform well for graph representation learning. In this paper, we solve this mystery by presenting Graphormer, which is built upon the standard Transformer architecture, and could attain excellent results on a broad range of graph representation learning tasks, especially on the recent OGB Large-Scale Challenge. Our key insight to utilizing Transformer in the graph is the necessity of effectively encoding the structural information of a graph into the model. To this end, we propose several simple yet effective structural encoding methods to help Graphormer better model graph-structured data. Besides, we mathematically characterize the expressive power of Graphormer and exhibit that with our ways of encoding the structural information of graphs, many popular GNN variants could be covered as the special cases of Graphormer.

Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, Tie-Yan Liu• 2021

Related benchmarks

TaskDatasetResultRank
Graph ClassificationPROTEINS
Accuracy78.98
1383
Graph ClassificationMUTAG
Accuracy90.99
1229
Node ClassificationCora
Accuracy86.4
1225
Node ClassificationCiteseer
Accuracy74.6
1037
Node ClassificationCiteseer (test)
Accuracy0.3318
1013
Node ClassificationChameleon
Accuracy53.8
936
Node ClassificationCornell
Accuracy70.59
900
Node ClassificationWisconsin
Accuracy84
898
Node ClassificationTexas
Accuracy0.767
859
Node ClassificationSquirrel
Accuracy34.6
815
Showing 10 of 204 rows
...

Other info

Code

Follow for update