Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Forgive or forget: Understanding the context of hate in audio retrieval systems

About

Handling toxic retrieval in text-to-audio systems is challenging due to contextual dependencies. Existing strategies (e.g., rephrasing, summarization) risk altering intent or omitting details. We propose a post hoc causal debiasing framework with a sentiment-controlled mediator to preserve semantic relevance while suppressing harmful speech. Our approach is model-agnostic and integrates seamlessly with existing retrieval pipelines. We introduce two variants: Forgive, which re-ranks and filters toxic audio via logit adjustment, and Forget, which generates counterfactual toxic prompts to mitigate harmful retrievals. Experiments show consistent toxicity reduction with minimal loss in retrieval accuracy, improving both safety and reliability.

Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Shekhar Nayak• 2026

Related benchmarks

TaskDatasetResultRank
Text-to-Audio RetrievalClotho V1 (test)
SR98
24
Showing 1 of 1 rows

Other info

Follow for update