Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SpecBench

Benchmarks

Task NameDataset NameSOTA ResultTrend
Speculative DecodingSpecBench
AVG SR900.7
47
Specification AlignmentSPECBENCH Average over scenarios
Safety Score93.73
33
Speculative DecodingSpecBench v1 (test)
WRIT τ4.9
12
Speculative DecodingSpecBench Qwen-2-7B-Instruct (test)
Overall Mean Score3.65
5
Preference ElicitationSpecBench Task 1
Improvement Rate (T0->T5)74
4
Relational Data SynthesisSpecBench Curated Domains Cold-start spec-mode
CSC1
4
Specification AlignmentSPECBENCH
Safety Score96.8
4
Speculative DecodingSpecBench Vicuna-33B v1.3 (test)
MT Score3.52
4
LLM GenerationSpecBench
Tokens/s116.95
3
Showing 9 of 9 rows