Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

RAR

Benchmarks

Task NameDataset NameSOTA ResultTrend
Provenance TracingRAR
TPR@1%FPR100
33
Downstream retrievalRAR-B
ARC nDCG@516.2
24
Image Attribution DetectionRAR generated images
AUC100
20
Dataset usage estimationRAR-XXL
MAE0.041
12
Dataset usage estimationRAR-XL
MAE4.5
12
Autoregressive Visual WatermarkingRAR-XL generation
Fidelity Score (Baseline)1
10
Image AttributionRAR
NM/GM100
8
Medical ReasoningRaR Medicine
WR vs Base57.6
8
Scientific ReasoningRAR-SCIENCE (test)
Accuracy69.9
6
Medical Question AnsweringRaR-Medicine (test)
Length1,395
5
Membership Inference AttackRAR-XXL
TPR@1%FPR8.4
4
AttributionRAR
Model Pre-training Time (hours)20,000
4
Model Generation IdentificationRAR
TPR@5%FPR (NM/G)100
4
Identifying memorized samplesRAR-XXL memorized 1.0 (169 train samples)
AUC97.5
4
Sample AttributionRAR model derivative setting
NM/GM Score100
4
Member vs Generated InferenceRAR
TPR@1%FPR (NM vs G)99.9
4
Pairwise Preference EvaluationRaR Medicine
Pairwise Win Rate60.6
4
Showing 17 of 17 rows