Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Hallucination Detection on LLM-generated screenplays Story 3
Loading...
75
Precision
Atlas
68.968
70.534
72.1
73.666
Jul 1, 2026
Precision
Recall
F1 Score
Gold Count
TP
FP
FN
Updated 23d ago
Evaluation Results
Method
Method
Links
Precision
Recall
F1 Score
Gold Count
TP
FP
FN
Atlas
2026.07
75
81.8
78.3
-
-
-
-
Atlas
Evaluator=GPT-5.4
2026.07
75
81.8
78.3
11
9
3
2
LLM-as-a-Judge
2026.07
69.2
81.8
75
-
-
-
-
LLM-as-a-Judge
Evaluator=GPT-5.4
2026.07
69.2
81.8
75
11
9
4
2
Feedback
Search any
task
Search any
task