AffordanceLLM: Grounding Affordance from Vision Language Models

About

Affordance grounding refers to the task of finding the area of an object with which one can interact. It is a fundamental but challenging task, as a successful solution requires the comprehensive understanding of a scene in multiple aspects including detection, localization, and recognition of objects with their parts, of geo-spatial configuration/layout of the scene, of 3D shapes and physics, as well as of the functionality and potential interaction of the objects and humans. Much of the knowledge is hidden and beyond the image content with the supervised labels from a limited training set. In this paper, we make an attempt to improve the generalization capability of the current affordance grounding by taking the advantage of the rich world, abstract, and human-object-interaction knowledge from pretrained large-scale vision language models. Under the AGD20K benchmark, our proposed model demonstrates a significant performance gain over the competing methods for in-the-wild object affordance grounding. We further demonstrate it can ground affordance for objects from random Internet images, even if both objects and actions are unseen during training. Project site: https://jasonqsy.github.io/AffordanceLLM/

Shengyi Qian, Weifeng Chen, Min Bai, Xiong Zhou, Zhuowen Tu, Li Erran Li• 2024

Related benchmarks

Task	Dataset	Result
Affordance prediction	AGD20K unseen	KLD1.463	26
Affordance Grounding	ReasonAff (test)	gIoU48.49	21
Affordance Grounding	UMD (test)	gIoU43.11	18
Affordance Reasoning	ReasonAff	gIoU48.49	15
2D affordance grounding	AGD20K unseen (test)	KLD1.463	14
Affordance prediction	ReasonAff (test)	gIoU48.49	13
Affordance prediction	UMD	gIoU43.11	10
Affordance Reasoning	UMD Part Affordance (test)	gIoU43.11	8
Affordance Estimation	AGD20K zero-shot	KLD1.463	6

Showing 9 of 9 rows

Other info

Follow for update

@wizwand_team Discord