Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Narrative Generation on vessel trip descriptions (test)
Loading...
4.91
Relevance
openai/gpt-oss-120b
3.8596
4.1323
4.405
4.6777
Mar 8, 2026
Relevance
Faithfulness
Correctness
Updated 4mo ago
Evaluation Results
Method
Method
Links
Relevance
Faithfulness
Correctness
openai/gpt-oss-120b
LLM Model=openai/gpt-o...
2026.03
4.91
4.96
4.69
openai/gpt-oss-20b
LLM Model=openai/gpt-o...
2026.03
4.69
4.81
4.35
llama-3.3-70b-versatile
LLM Model=llama-3.3-70...
2026.03
4.47
4.77
4.36
qwen/qwen3-32b
LLM Model=qwen/qwen3-32b
2026.03
4.42
4.48
3.82
llama-3.1-8b-instant
LLM Model=llama-3.1-8b...
2026.03
3.9
3.84
3.27
Feedback
Search any
task
Search any
task