Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

L3

Benchmarks

Task NameDataset NameSOTA ResultTrend
Two-Sample TestingL3 d=512
Testing Power17.6
8
Two-Sample TestingL3 d=256
Testing Power28
8
Action-component payoff optimizationL3 warmup (off-diagonal)
Per-Interaction Payoff4.17
8
Literature-to-ProductionL3 Literature-to-Production
LPR86.7
4
Semantic ReasoningL3 (out-of-distribution)
Accuracy (Series Comparison)67
2
Showing 5 of 5 rows