Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
LLM-as-a-Judge Quality Evaluation on 92 Queries Qwen 80B Evaluation 1.0 (test)
Loading...
7.83
Comprehensiveness
NotebookLM
7.362
7.4835
7.605
7.7265
Apr 20, 2026
Comprehensiveness
Relevance
Coverage
Coherence
Appropriateness
Grammatical Correctness
Adherence to Constraints
Causal Reasoning
Safety/Bias
Average Score
Updated 3mo ago
Evaluation Results
Method
Method
Links
Comprehensiveness
Relevance
Coverage
Coherence
Appropriateness
Grammatical Correctness
Adherence to Constraints
Causal Reasoning
Safety/Bias
Average Score
NotebookLM
Tier=Pro
2026.04
7.83
8.1
8.1
9.12
8.23
9.74
8.35
8.33
9.88
8.7
AVA AI
2026.04
7.58
8.07
8.07
9.48
8.74
9.92
8.01
8.74
9.96
8.81
Perplexity AI
Version=Enterprise, We...
2026.04
7.38
8.34
8.34
9.28
8.83
9.67
8.35
8.28
10
8.77
Feedback
Search any
task
Search any
task