Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Dolly

Benchmarks

Task NameDataset NameSOTA ResultTrend
Instruction FollowingDolly
Rouge-L27.74
50
Watermark Detectiondolly_cw
Accuracy100
48
Instruction FollowingDolly Eval (test)
ROUGE-L29.69
42
Question AnsweringDolly Closed QA
ASR100
36
Instruction FollowingDolly
Score71.3
36
Hallucination detectionDolly AC (test)
AUC81.59
33
General Question Answering & Instruction FollowingDolly
MP Score22
24
Data Leakage AttackDolly
AP (alpha=0.5)98.5
24
MMLU EvaluationDolly
Accuracy30.94
24
Detection Accuracydolly_cw
Accuracy99.27
24
Instruction FollowingDolly
SBERT Similarity71.4
24
Instruction TuningDolly-15K alpha=5.0
Rouge-L35.79
22
Instruction TuningDolly-15K alpha=0.5
Rouge-L35.48
22
Instruction-tuningDolly
RougeL35.34
21
Hallucination DetectionDolly Llama2-13B (test)
Accuracy75.76
21
Hallucination DetectionDolly Llama2-7B (test)
Acc77.78
21
Scrubbing AttackDolly
AUC80
20
Hallucination DetectionDolly AC LLaMA3-8B
Recall83.92
19
Hallucination DetectionDolly AC LLaMA2-13B
Recall0.9741
19
Hallucination DetectionDolly AC LLaMA2-7B
Recall87.28
19
Instruction FollowingDolly Eval
A Win Count62
19
Open-ended generationDolly
Skywork Reward V2 Score0.961
18
Spoofing Attack DetectionDolly CW
WCS8.88
18
Instruction Following EvaluationDolly Out-of-Distribution
GPT-4o Score49.9
17
Language GenerationDolly databricks 15k (test)
ROUGE-L29.7
14
Showing 25 of 50 rows