Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data Attribution

About

Scalable data attribution methods typically assign isolated utility scores to individual training examples. This prevalent additive assumption fundamentally fails to capture critical subset dynamics, including data redundancy and complementary coverage. In this work, we reframe attribution as subset-level counterfactual utility prediction and introduce GRASP, an interaction-aware surrogate. Grounded in a theoretical smoothness lower bound, GRASP explicitly models subset interactions through a quadratic geometric penalty. To achieve pretraining-scale efficiency without relying on hidden oracle tuning, we couple low-dimensional feature sketches with a strictly finite lower-confidence bound selection protocol. Extensive subset-retraining evaluations demonstrate that GRASP decisively outperforms existing scalable baselines. It more than doubles the task-level rank correlation for counterfactual subset fidelity while reducing upfront artifact construction costs by nearly an order of magnitude. Downstream diagnostics further show that this scoring mechanism transfers to language model curation and cross-domain vision selection, establishing a robust foundation for optimizing massive pretraining corpora.

Yue Min, Ruining Chen, Yujun Li• 2026

Related benchmarks

TaskDatasetResultRank
Noisy label detectionCIFAR-10 1k subset 10% label noise
AUC (%)80.32
8
Vision data selectionCIFAR-10 1k 10% label noise
Accuracy (20% Selected Data)56.1
8
Malicious or compromised fine-tuning example detectionMisleading information and safety risks dataset
F1 Score30
6
Related-domain retrievalCC-10B WebOrganizer Topic labels (candidate pool)
Mean Normalized Recall@10007.121
6
Related-domain retrievalCC-10B WebOrganizer Format labels (candidate pool)
Mean Normalized Recall@10004.1
6
Counterfactual Subset-Utility EvaluationBasicSkills Common Knowledge
Spearman Correlation (ρ)0.312
4
Counterfactual Subset-Utility EvaluationBasicSkills Logical Reasoning
Spearman Correlation ($ ho$)0.457
4
Counterfactual Subset-Utility EvaluationOpenBookQA
Spearman Correlation0.35
4
Counterfactual Subset-Utility EvaluationCommonsenseQA
Spearman Correlation ($ ho$)0.26
4
Counterfactual Subset-Utility EvaluationSciQ
Spearman Correlation (ρ)0.424
4
Showing 10 of 12 rows

Other info

Follow for update