Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CN

Benchmarks

Task NameDataset NameSOTA ResultTrend
Multi-Agent Reinforcement LearningCN rac-dist
Mean Episodic Reward888
21
Multi-Agent Reinforcement LearningCN rdist
Mean Episodic Reward-161
21
Multi-Agent Reinforcement LearningCN rdete
Mean Episodic Reward-154
21
Role-Playing Evaluation (Conversational-Naturalness)CN
Win Rate65
9
CN task streamCN (Medium)
Backward Transfer35
8
CN task streamCN (Expert)
Backward Transfer5.83
8
Multi-agent Continual CooperationCN Medium
Forward Transfer16
7
Multi-agent Continual CooperationCN Expert
Forward Transfer6.74
7
Cooperative NavigationCN MPE hard
Mean Episode Reward3.37
7
Cooperative NavigationCN MPE medium
Mean Episode Reward3.21
7
Soft Query AnsweringCN15k
1P Score16.6
6
Backdoor AttackCN (test)
Runtime (s)28.3
4
Intent PredictionCN
Accuracy55.2
4
Function InvocationCN Ver. Dual
Token Usage1,377.9
3
Function InvocationCN (Single)
Invocation Accuracy0.89
3
Competitive ratio of ε-LDP online (Wε) vs. ε-LDP offline (Lε) stoppingCn independent non-negative random variables
Lower Bound Competitive Ratio1
1
Competitive ratio of ε-LDP online (Wε) vs. non-private offline (M) stoppingCn independent non-negative random variables
Lower Bound Competitive Ratio2
1
Competitive ratio of ε-LDP online (Wε) vs. non-private online (V) stoppingCn independent non-negative random variables
Lower Bound Competitive Ratio1
1
Competitive ratio of non-private online (V) vs. non-private offline (M) stoppingCn independent non-negative random variables
Lower Bound Competitive Ratio1
1
Showing 19 of 19 rows