Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Hallucination Detection on LLM-generated screenplays Story 2
Loading...
100
Precision
Atlas
83.984
88.142
92.3
96.458
Jul 1, 2026
Precision
Recall
F1 Score
Gold Count
True Positives (TP)
False Positives (FP)
False Negatives (FN)
Updated 23d ago
Evaluation Results
Method
Method
Links
Precision
Recall
F1 Score
Gold Count
True Positives (TP)
False Positives (FP)
False Negatives (FN)
Atlas
2026.07
100
84.6
91.7
-
-
-
-
Atlas
Evaluator=GPT-5.4
2026.07
100
84.6
91.7
13
11
0
2
LLM-as-a-Judge
2026.07
84.6
84.6
84.6
-
-
-
-
LLM-as-a-Judge
Evaluator=GPT-5.4
2026.07
84.6
84.6
84.6
13
11
2
2
Feedback
Search any
task
Search any
task