PPI++: Efficient Prediction-Powered Inference
About
We present PPI++: a computationally lightweight methodology for estimation and inference based on a small labeled dataset and a typically much larger dataset of machine-learning predictions. The methods automatically adapt to the quality of available predictions, yielding easy-to-compute confidence sets -- for parameters of any dimensionality -- that always improve on classical intervals using only the labeled data. PPI++ builds on prediction-powered inference (PPI), which targets the same problem setting, improving its computational and statistical efficiency. Real and synthetic experiments demonstrate the benefits of the proposed adaptations.
Anastasios N. Angelopoulos, John C. Duchi, Tijana Zrnic• 2023
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Population property estimation | DICES | Bias (MAE)0.06 | 92 | |
| LLM evaluation human preference | PPE Human Preference track | MSE / PPI0.283 | 28 | |
| LLM evaluation correctness | PPE Correctness track | MSE / PPI0.276 | 20 | |
| Regression coefficient estimation | Politeness dataset 5512 online requests (full) | Required Labeled Observations697 | 17 | |
| Log-income regression (coefficient estimation) | US Census Data (whole population) | Required Labeled Observations1.03e+4 | 17 | |
| Bias Reduction Estimation | Private Healthcare Census Setting | Average MAPE Difference-15.22 | 15 | |
| Mean Estimation | CivilComments-WILDS (test) | Required Labeled Samples1.39e+3 | 14 | |
| LLM win-rate estimation ranking | LLM benchmark (Appendix) | Spearman Correlation1 | 14 | |
| Regression coefficient estimation | WineEnthusiast | Labeled Samples2.63e+3 | 12 | |
| Confidence interval estimation | NIH ChestX-ray14 (test) | Required Labeled Observations4.49e+3 | 12 |
Showing 10 of 28 rows