Osprey: Pixel Understanding with Visual Instruction Tuning

About

Multimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning. However, current MLLMs primarily focus on image-level or box-level understanding, falling short in achieving fine-grained vision-language alignment at pixel level. Besides, the lack of mask-based instruction data limits their advancements. In this paper, we propose Osprey, a mask-text instruction tuning approach, to extend MLLMs by incorporating fine-grained mask regions into language instruction, aiming at achieving pixel-wise visual understanding. To achieve this goal, we first meticulously curate a mask-based region-text dataset with 724K samples, and then design a vision-language model by injecting pixel-level representation into LLM. Specifically, Osprey adopts a convolutional CLIP backbone as the vision encoder and employs a mask-aware visual extractor to extract precise visual mask features from high resolution input. Experimental results demonstrate Osprey's superiority in various region understanding tasks, showcasing its new capability for pixel-level instruction tuning. In particular, Osprey can be integrated with Segment Anything Model (SAM) seamlessly to obtain multi-granularity semantics. The source code, dataset and demo can be found at https://github.com/CircleRadon/Osprey.

Yuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, Jianke Zhu• 2023

Related benchmarks

Task	Dataset	Result
Semantic segmentation	Cityscapes (val)	mIoU49.78	527
Object Hallucination	POPE Popular	F1 Score87.5	372
Object Hallucination	POPE Adversarial	Accuracy85.33	353
Object Hallucination	POPE (Random)	F1 Score88.97	324
Panoptic Segmentation	Cityscapes (val)	PQ50.64	288
Panoptic Segmentation	ADE20K 150 59 (val)	Panoptic Quality (PQ)41.89	35
Referring expression generation	RefCOCOg (val)	METEOR16.6	31
Instance Segmentation	ADE20K 150 59 (val)	AP41.24	30
Video Referring	VideoRefer-Bench-D	SC3.3	23
Video Understanding	VideoRefer-BenchQ	Overall Accuracy39.9	22

Showing 10 of 37 rows

Other info

Code

Follow for update

@wizwand_team Discord