Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

GPT4-LLM

Benchmarks

Task NameDataset NameSOTA ResultTrend
Over-refusal evaluationPoisoned GPT4-LLM subset
Informative Refusal Rate35.7
27
Showing 1 of 1 rows