Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Response Quality Evaluation on ID in-domain (150 samples)
Loading...
134
Count
GMM
11.28
43.14
75
106.86
Aug 4, 2025
Count
Avg LLM Relevance
Avg Human Relevance
Avg LLM Correctness
Avg Human Correctness
Updated 19d ago
Evaluation Results
Method
Method
Links
Count
Avg LLM Relevance
Avg Human Relevance
Avg LLM Correctness
Avg Human Correctness
GMM
Classification Outcome...
2025.08
134
4.66
4.75
4.82
4.39
GPT-4o
Classification Outcome...
2025.08
126
4.69
4.73
4.81
4.35
GPT-4o
Classification Outcome...
2025.08
24
4.17
4.56
4.46
4.1
GMM
Classification Outcome...
2025.08
16
4.19
4.34
4.19
3.75
Feedback
Search any
task
Search any
task