AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

About

As Large Language Models (LLMs) and generative AI become more widespread, the content safety risks associated with their use also increase. We find a notable deficiency in high-quality content safety datasets and benchmarks that comprehensively cover a wide range of critical safety areas. To address this, we define a broad content safety risk taxonomy, comprising 13 critical risk and 9 sparse risk categories. Additionally, we curate AEGISSAFETYDATASET, a new dataset of approximately 26, 000 human-LLM interaction instances, complete with human annotations adhering to the taxonomy. We plan to release this dataset to the community to further research and to help benchmark LLM models for safety. To demonstrate the effectiveness of the dataset, we instruction-tune multiple LLM-based safety models. We show that our models (named AEGISSAFETYEXPERTS), not only surpass or perform competitively with the state-of-the-art LLM-based safety models and general purpose LLMs, but also exhibit robustness across multiple jail-break attack categories. We also show how using AEGISSAFETYDATASET during the LLM alignment phase does not negatively impact the performance of the aligned models on MT Bench scores. Furthermore, we propose AEGIS, a novel application of a no-regret online adaptation framework with strong theoretical guarantees, to perform content moderation with an ensemble of LLM content safety experts in deployment

Shaona Ghosh, Prasoon Varshney, Erick Galinkin, Christopher Parisien• 2024

Related benchmarks

Task	Dataset	Result
Response Harmfulness Detection	HarmBench	F1 Score77.7	100
Response Harmfulness Detection	XSTEST-RESP	Response Harmfulness F160.4	76
Response Harmfulness Detection	Beavertails	F1 Score74.7	59
Safety Classification	SafeRLHF	F1 Score0.593	48
Harmfulness Detection	WildGuard	Macro F1 Score78.5	47
Safety Classification	WildGuardMix (test)	F1 Score82.1	47
Toxicity Detection	ToxicChat	F1 Score0.73	45
Harmfulness Detection	OpenAI Moderation	Macro F1 Score74.7	45
Prompt Harmfulness Detection	AegisSafety (test)	F1 Score84.8	41
Response Harmfulness Detection	SafeRLHF	F1 Score59.3	41

Showing 10 of 34 rows

Other info

Follow for update

@wizwand_team Discord