Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LaSagnA: Language-based Segmentation Assistant for Complex Queries

About

Recent advancements have empowered Large Language Models for Vision (vLLMs) to generate detailed perceptual outcomes, including bounding boxes and masks. Nonetheless, there are two constraints that restrict the further application of these vLLMs: the incapability of handling multiple targets per query and the failure to identify the absence of query objects in the image. In this study, we acknowledge that the main cause of these problems is the insufficient complexity of training queries. Consequently, we define the general sequence format for complex queries. Then we incorporate a semantic segmentation task in the current pipeline to fulfill the requirements of training data. Furthermore, we present three novel strategies to effectively handle the challenges arising from the direct integration of the proposed format. The effectiveness of our model in processing complex queries is validated by the comparable results with conventional methods on both close-set and open-set semantic segmentation datasets. Additionally, we outperform a series of vLLMs in reasoning and referring segmentation, showcasing our model's remarkable capabilities. We release the code at https://github.com/congvvc/LaSagnA.

Cong Wei, Haoxian Tan, Yujie Zhong, Yujiu Yang, Lin Ma• 2024

Related benchmarks

TaskDatasetResultRank
Referring Image SegmentationRefCOCO (val)--
259
Referring Expression SegmentationRefCOCO (testA)
cIoU78.7
257
Referring Image SegmentationRefCOCO+ (test-B)--
252
Diagram UnderstandingAI2D
Accuracy0.00e+0
247
Referring Expression SegmentationRefCOCO+ (testA)
cIoU70.6
230
Referring Image SegmentationRefCOCO (test A)--
230
Referring Expression SegmentationRefCOCO+ (val)
cIoU66.4
223
Referring Expression SegmentationRefCOCO (testB)
cIoU73.8
213
Referring Expression SegmentationRefCOCO (val)
cIoU76.8
212
Referring Expression SegmentationRefCOCO+ (testB)
cIoU60.1
210
Showing 10 of 32 rows

Other info

Follow for update