Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Helpfulness Evaluation on Alpaca
Loading...
90
Helpfulness Score
No Defense
-2.56
21.47
45.5
69.53
Aug 9, 2025
Helpfulness Score
Updated 18d ago
Evaluation Results
Method
Method
Links
Helpfulness Score
No Defense
Base LLM=ChatGPT
2025.08
90
Self-Reminder
Base LLM=ChatGPT
2025.08
90
Self-Examination
Base LLM=ChatGPT
2025.08
90
ICD
Base LLM=ChatGPT
2025.08
88
Context Filtering
Base LLM=ChatGPT
2025.08
88
No Defense
Base LLM=Llama2
2025.08
62
Context Filtering
Base LLM=Llama2
2025.08
60
No Defense
Base LLM=Vicuna
2025.08
59
Context Filtering
Base LLM=Vicuna
2025.08
57
Self-Reminder
Base LLM=Vicuna
2025.08
56
Self-Examination
Base LLM=Vicuna
2025.08
56
Self-Reminder
Base LLM=Llama2
2025.08
55
Safe Decoding
Base LLM=Llama2
2025.08
52
ICD
Base LLM=Vicuna
2025.08
51
Safe Decoding
Base LLM=Vicuna
2025.08
50
Intention Analysis
Base LLM=Vicuna
2025.08
33
ICD
Base LLM=Llama2
2025.08
21
Self-Examination
Base LLM=Llama2
2025.08
5
Intention Analysis
Base LLM=ChatGPT
2025.08
4
Intention Analysis
Base LLM=Llama2
2025.08
1
Feedback
Search any
task
Search any
task