Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Prompts

Benchmarks

Task NameDataset NameSOTA ResultTrend
Text-to-Image Generationprompts 500 sampled
HPSv20.3345
36
Speculative Decoding20 Prompts across 4 Task Categories
Mean Expected Tokens per Speculation Step6.55
20
Jailbreak Prompt Quality Evaluation500 randomly sampled prompts
Similarity81
16
Distribution-distance evaluationPrompts 100 (evaluation)
Distinct-N (WM)94.1
14
Creative Plot Generation160 prompts NQD (test)
Character Development8.67
13
Over-generation attack1000 prompts (test)
Succ. @≥ 188.2
8
Text-to-Image Generationprompts 10 randomly sampled
Inference Time (s)2.2322
6
Property-based retrievalPrompts (test)
MAP0.48
6
Platform identification1000 prompts (held-out)
CPRSD l(f)1.1
5
Text-to-Video1,024 prompts (held-out)
VQ4.81
5
Panorama Generation14 prompts 1000 panoramas of dimensions 512x4608
Intra-LPIPS0.58
4
Text-to-Image Generation400 prompts (test)
HPSv229.0533
4
Human Preference Evaluation (Harmlessness)1,172 Prompts (test)
Win Count (CS)677
3
Human Preference Evaluation (Helpfulness)1,172 prompts (test)
CS Wins695
3
Steering LLM states50 prompts
LogFreq (d)1.6666
3
3D Scene Editing15 distinct single-task prompts
LLM Time10.63
3
LLM agent alignment evaluation1000 prompts (test)
Usefulness Score1
2
Word Count Adherence180 held-out prompts Length Far OOD 8–12k
Length Adherence Ratio72
1
Word Count Adherence180 held-out prompts Length Near OOD 4–8k
Length Adherence Ratio87
1
Word Count Adherence180 Prompts Length ID 1–4k (test)
Length Adherence Ratio99
1
Story Generation Quality180 held-out prompts Far OOD 8–12k
Story Quality Score44.1
1
Story Generation Quality180 prompts Near OOD 4–8k (held-out)
Story Quality Score48.2
1
Showing 22 of 22 rows