Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
LLM-as-a-Judge Quality Evaluation on 92 Queries Grok 3 Evaluation 1.0 (test)
Loading...
7.53
Comprehensiveness
AVA AI
6.2716
6.5983
6.925
7.2517
Apr 20, 2026
Comprehensiveness
Relevance
Coverage
Coherence
Appropriateness
Grammatical Correctness
Adherence to Constraints
Causal Reasoning
Safety/Bias
Average Score
Updated 3mo ago
Evaluation Results
Method
Method
Links
Comprehensiveness
Relevance
Coverage
Coherence
Appropriateness
Grammatical Correctness
Adherence to Constraints
Causal Reasoning
Safety/Bias
Average Score
AVA AI
2026.04
7.53
7.53
7.53
8.48
7.99
9.8
8.22
7.43
9.8
8.35
NotebookLM
Tier=Pro
2026.04
6.82
6.69
6.69
7.77
7.03
9.51
7.58
6.63
9.64
7.71
Perplexity AI
Version=Enterprise, We...
2026.04
6.32
7.08
7.08
8.1
7.45
9.37
7.82
6.5
9.51
7.77
Feedback
Search any
task
Search any
task