Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Conversation Simulation on Humanual-Chat
Loading...
28.2
Primary Quality Score
GPT 5.5
6.984
12.492
18
23.508
Jun 12, 2026
Primary Quality Score
Updated 1mo ago
Evaluation Results
Method
Method
Links
Primary Quality Score
GPT 5.5
2026.06
28.2
OSIM 8B
Model=OSIM, Parameter...
2026.06
28.2
Others *
2026.06
25.8
Qwen3 8B Inst
Model=Qwen3-8B-Instruc...
2026.06
24.7
Claude Opus 4.7
2026.06
22.6
Qwen 3.6 Plus
2026.06
22.2
Gemini 3.1 Pro
2026.06
21
Base 8B
Model=Qwen3-8B-Base, P...
2026.06
12
OSIM 8B-Mid
Model=OSIM, Parameter...
2026.06
7.8
Feedback
Search any
task
Search any
task