Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

VecEval

Benchmarks

Task NameDataset NameSOTA ResultTrend
Explanation GenerationVecEval (val)
B Score93.26
9
Context-aware EvaluationVecEval (val)
Lingo-Judge Accuracy78.61
9
Urgency PredictionVecEval (val)
Intervention Failure Rate12.12
9
Showing 3 of 3 rows