Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Detection on Algospeak Abbreviation 1.0 (test)
Loading...
0.98
Adjusted R2
GPT-4o-m
-0.0912
0.1869
0.465
0.7431
May 7, 2026
Adjusted R2
Spearman Correlation Significance
Majority Fit Estimation
Significance Count
Updated 26d ago
Evaluation Results
Method
Method
Links
Adjusted R2
Spearman Correlation Significance
Majority Fit Estimation
Significance Count
GPT-4o-m
2026.05
0.98
-
-
-
Grok
2026.05
0.98
-
-
-
Llama
2026.05
0.96
-
-
-
GPT-4o
2026.05
0.95
-
-
-
Claude
2026.05
0.92
-
-
-
Qwen
2026.05
0.91
-
-
-
Mistral
2026.05
-0.05
-
-
-
Significance Count
2026.05
-
-
-
7
Feedback
Search any
task
Search any
task