Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

benchmark suite

Benchmarks

Task NameDataset NameSOTA ResultTrend
General Language Understanding10-benchmark suite average
Average Accuracy73.81
15
Generalization across multiple tasksFull Benchmark Suite Aggregate
Average Accuracy82.2
10
Offline Policy AdaptationFull Benchmark Suite Gravity and Morph shifts
Total Score820.8
7
Visual instruction tuningBenchmark Suite Aggregate
Average Score65.1
6
Quantum architecture compilation36-circuit benchmark suite Combined
Geometric-mean Execution Time19,507.5
4
Semantic Segmentation20-domain benchmark suite
h-mean58.05
4
Text Generationbenchmark suite 20-task
Average End-to-End Latency (ms)2,520
3
Unique bug detectionFull benchmark suite Very Large
TP6
3
LLM Pruning EvaluationFull Benchmark Suite
S_tot Retention75.15
2
Showing 9 of 9 rows