Trajectory balance: Improved credit assignment in GFlowNets

About

Generative flow networks (GFlowNets) are a method for learning a stochastic policy for generating compositional objects, such as graphs or strings, from a given unnormalized density by sequences of actions, where many possible action sequences may lead to the same object. We find previously proposed learning objectives for GFlowNets, flow matching and detailed balance, which are analogous to temporal difference learning, to be prone to inefficient credit propagation across long action sequences. We thus propose a new learning objective for GFlowNets, trajectory balance, as a more efficient alternative to previously used objectives. We prove that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution. In experiments on four distinct domains, we empirically demonstrate the benefits of the trajectory balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequences and large action spaces.

Nikolay Malkin, Moksh Jain, Emmanuel Bengio, Chen Sun, Yoshua Bengio• 2022

Related benchmarks

Task	Dataset	Result
Molecule Design	Molecule Design (test)	Mode Coverage (R>7.5)2.51e+3	26
Anti-Microbial Peptide Generation	AMP generation	Top 100 Reward0.85	5
Probabilistic Modeling	HyperGrid Original	Final L1 Distance0.438	4
SMILES generation	SMILES (val)	Accuracy99.8	4
Probabilistic Modeling	Multiplicative Coprime HyperGrid	Final L1 Distance0.137	4
Probabilistic Modeling	Bitwise XOR HyperGrid	Final L1 Distance1.74	4
Probabilistic Modeling	Cosine HyperGrid	Final L1 Distance0.449	4
Molecule Generation	Molecule Design	Modes (R > 7.5)1.92e+3	3

Showing 8 of 8 rows

Other info

Code

Follow for update

@wizwand_team Discord