Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ChatDoctor

Benchmarks

Task NameDataset NameSOTA ResultTrend
Untargeted Reconstruction AttackChatDoctor 250 adversarial prompts
Repeat Prompt40
42
Answer AccuracyChatDoctor
BRT Accuracy40.8
26
Retrieval-based Question AnsweringChatDoctor
ROUGE-L13.6
23
Privacy Leakage and Utility AssessmentChatDoctor iCliniq
Retrieved Context Count750
10
TracingChatDoctor (test)
TSR99.75
10
Instruction-followingChatDoctor
BRT46.9
9
Factual Consistency EvaluationChatDoctor
FC Score4.08
7
Medical Question AnsweringChatDoctor
Task Performance64.89
7
Preprocessing efficiencyChatDoctor
Preprocessing Time (s/sample)13.52
6
Property Inference DefenseChatDoctor QA
Generation Attack MAE0.13
6
Property Inference DefenseChatDoctor CC
Generation Attack MAE0.0354
6
QAChatDoctor
MAE0.0396
5
CCChatDoctor
MAE0.0304
5
Privacy-Utility Trade-off EvaluationChatDoctor worst-case attack setting
CRR97.4
5
Knowledge Base ExtractionChatDoctor RAG-Thief attack
CRR73.4
5
Knowledge Base ExtractionChatDoctor Pirates attack
CRR97.4
5
Monitor AccuracyChatDoctor
Monitor Accuracy96
4
Privacy extraction attackChatDoctor
Privacy Leakage Rate17
4
Backdoor DetectionChatDoctor
F1-score0.0058
3
Showing 19 of 19 rows