Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

About

Augmenting large language models (LLMs) to understand audio -- including non-speech sounds and non-verbal speech -- is critically important for diverse real-world applications of LLMs. In this paper, we propose Audio Flamingo, a novel audio language model with 1) strong audio understanding abilities, 2) the ability to quickly adapt to unseen tasks via in-context learning and retrieval, and 3) strong multi-turn dialogue abilities. We introduce a series of training techniques, architecture design, and data strategies to enhance our model with these abilities. Extensive evaluations across various audio understanding tasks confirm the efficacy of our method, setting new state-of-the-art benchmarks. Our demo website is https://audioflamingo.github.io/ and the code is open-sourced at https://github.com/NVIDIA/audio-flamingo.

Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, Bryan Catanzaro• 2024

Related benchmarks

Task	Dataset	Result
Audio Captioning	AudioCaps (test)	CIDEr0.846	222
Musical Instrument Classification	NSynth	Accuracy77.1	123
Audio Captioning	Clotho 2.1 (test)	--	75
Audio Captioning	AudioCaps	CIDEr54.6	66
Audio Reasoning	MMAR (test)	Average Score26.6	57
Audio Classification	US8K (test)	R@1 Accuracy0.75	56
Multiple-choice audio understanding	MMAU mini (test)	Average Accuracy16.6	39
Emotion Recognition	RAVDESS (test)	Accuracy0.209	29
Classification	GTZAN (test)	Accuracy67.9	23
Speech-to-speech translation	CVSS-C + SpeechMatrix FR-EN (test)	BLEU21.53	20

Showing 10 of 27 rows

Other info

Code

Follow for update

@wizwand_team Discord