Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Hallucination Detection on LLM-generated screenplays Story 1
Loading...
100
Precision
Atlas
95
97.5
100
102.5
Jul 1, 2026
Precision
Recall
F1 Score
Gold Count
True Positives (TP)
False Positives (FP)
False Negatives (FN)
Updated 23d ago
Evaluation Results
Method
Method
Links
Precision
Recall
F1 Score
Gold Count
True Positives (TP)
False Positives (FP)
False Negatives (FN)
Atlas
2026.07
100
75
85.7
-
-
-
-
LLM-as-a-Judge
2026.07
100
68.8
81.5
-
-
-
-
Atlas
Evaluator=GPT-5.4
2026.07
100
75
85.7
16
12
0
4
LLM-as-a-Judge
Evaluator=GPT-5.4
2026.07
100
68.8
81.5
16
11
0
5
Feedback
Search any
task
Search any
task