Multi-Hop Knowledge Graph Reasoning with Reward Shaping

About

Multi-hop reasoning is an effective approach for query answering (QA) over incomplete knowledge graphs (KGs). The problem can be formulated in a reinforcement learning (RL) setup, where a policy-based agent sequentially extends its inference path until it reaches a target. However, in an incomplete KG environment, the agent receives low-quality rewards corrupted by false negatives in the training data, which harms generalization at test time. Furthermore, since no golden action sequence is used for training, the agent can be misled by spurious search trajectories that incidentally lead to the correct answer. We propose two modeling advances to address both issues: (1) we reduce the impact of false negative supervision by adopting a pretrained one-hop embedding model to estimate the reward of unobserved facts; (2) we counter the sensitivity to spurious paths of on-policy RL by forcing the agent to explore a diverse set of paths using randomly generated edge masks. Our approach significantly improves over existing path-based KGQA models on several benchmark datasets and is comparable or better than embedding-based models.

Xi Victoria Lin, Richard Socher, Caiming Xiong• 2018

Related benchmarks

Task	Dataset	Result
Link Prediction	FB15k-237 (test)	Hits@1054.4	419
Link Prediction	WN18RR (test)	Hits@1054.2	380
Link Prediction	NELL-995 (test)	MRR23.1	27
Multi-hop Link Prediction	NELL-995 3-hop	MRR0.512	8
Multi-hop Link Prediction	FB15k-237 2-hop	MRR0.239	8
Multi-hop Link Prediction	NELL-995 2-hop	MRR57.3	8
Multi-hop Link Prediction	FB15k-237 3-hop	MRR0.191	7

Showing 7 of 7 rows

Other info

Follow for update

@wizwand_team Discord