Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Multi-task Language Evaluation on SEA-HELM
Loading...
83
Instruction Following (SEA-IFEval)
SiamGPT-32B
52.84
60.67
68.5
76.33
Dec 22, 2025
Instruction Following (SEA-IFEval)
Multi-Turn Dialogue (SEA-MTBench)
NLG (Translation + Summarization)
NLU (QA + Sentiment)
NLR (NLI + Causal Reasoning)
Safety Score
Average Score
Updated 5mo ago
Evaluation Results
Method
Method
Links
Instruction Following (SEA-IFEval)
Multi-Turn Dialogue (SEA-MTBench)
NLG (Translation + Summarization)
NLU (QA + Sentiment)
NLR (NLI + Causal Reasoning)
Safety Score
Average Score
SiamGPT-32B
Number of Parameters=3...
2025.12
83
75.81
42.06
67.95
68.59
44.19
63.6
Typhoon 2.5
Number of Parameters=3...
2025.12
79
76.16
56.7
65.56
55.54
29.68
60.44
OTG-R1
Number of Parameters=3...
2025.12
54
59.69
54.31
59.89
65.38
41.42
55.78
Feedback
Search any
task
Search any
task