Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Story

Benchmarks

Task NameDataset NameSOTA ResultTrend
Creative WritingStory
Semantic Diversity38.6
20
Story generationStory
Diversity8.36
19
Question AnsweringStory
Exact Match (EM)46.7
14
Single change-point detectionStory
WD0.207
12
Open-ended Text GenerationStory (test)
Diversity (DIV)0.96
12
Machine Text DetectionStory
Rewrite AUC (Claude 3.5)0.998
11
Multiple change-point detectionStory dataset GPT-5-mini K=5
WD0.44
6
Text GenerationStory
Coherence Win Rate63.6
4
Showing 8 of 8 rows