Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Energy Transformer

About

Our work combines aspects of three promising paradigms in machine learning, namely, attention mechanism, energy-based models, and associative memory. Attention is the power-house driving modern deep learning successes, but it lacks clear theoretical foundations. Energy-based models allow a principled approach to discriminative and generative tasks, but the design of the energy functional is not straightforward. At the same time, Dense Associative Memory models or Modern Hopfield Networks have a well-established theoretical foundation, and allow an intuitive design of the energy function. We propose a novel architecture, called the Energy Transformer (or ET for short), that uses a sequence of attention layers that are purposely designed to minimize a specifically engineered energy function, which is responsible for representing the relationships between the tokens. In this work, we introduce the theoretical foundations of ET, explore its empirical capabilities using the image completion task, and obtain strong quantitative results on the graph anomaly detection and graph classification tasks.

Benjamin Hoover, Yuchen Liang, Bao Pham, Rameswar Panda, Hendrik Strobelt, Duen Horng Chau, Mohammed J. Zaki, Dmitry Krotov• 2023

Related benchmarks

TaskDatasetResultRank
Graph ClassificationPROTEINS
Accuracy90.3
1383
Graph ClassificationMUTAG
Accuracy96.6
1229
Graph ClassificationNCI1
Accuracy90.1
707
Graph ClassificationENZYMES
Accuracy99.8
419
Graph ClassificationDD
Accuracy95.9
309
Graph ClassificationNCI109
Accuracy90.5
275
Graph ClassificationPROTEINS TUDataset
Accuracy90.3
44
Graph ClassificationNCI1 TUDataset
Accuracy90.1
44
Graph ClassificationMutagenicity
Accuracy98.7
35
Graph ClassificationMUTAGENICITY TUDataset
Accuracy98.7
31
Showing 10 of 30 rows

Other info

Code

Follow for update