| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Language Modeling | GPT-2 | Delta PPL0.44 | 32 | |
| Classification | GPT-2 Phase 2 (test) | Accuracy89.18 | 28 | |
| Time-to-accuracy Language Modeling | GPT-2 Medium (train) | Training Time (h)29.47 | 27 | |
| Language Modeling | GPT-2 Pre-training (val) | Validation Loss2.493 | 21 | |
| Language Modeling | GPT-2 Evaluation Set | Hyper-Prior BPT29.91 | 20 | |
| Privacy Attack | GPT-2 Small (124M) | Validation CE Loss0.68 | 12 | |
| Language Modeling | GPT-2 Pretraining Data (train) | Training Loss2.9167 | 12 | |
| Layer Pruning | GPT-2-Medium standardized evaluator | Delta (%)109.1 | 10 | |
| Linguistic Steganography | GPT-2 | Avg KLD (bits/token)0 | 10 | |
| Privacy-Preserving Transformer Inference | GPT-2 Transformer Block | Latency (s)0.2 | 10 | |
| Language Modeling | GPT-2 124M held-out (test) | Perplexity17.33 | 10 | |
| Language Modeling | GPT-2 Medium (val) | Runtime87.89 | 9 | |
| SAE Feature Attribution | GPT-2 Small 128K SAE | Input Attribution (%)32.16 | 9 | |
| Interpreting Sparse Key-Value Features | GPT-2 Small | Input Success Rate39.32 | 9 | |
| Machine-Generated Text Detection | GPT-2 (full) | Acc91.1 | 9 | |
| Data Extraction | GPT-2 (train) | Pearson r0.48 | 8 | |
| Private Embedding Lookup | GPT-2 | VecGen423.9253 | 6 | |
| Concentration of target information | GPT-2 Small suite Aggregate (test) | Gini Coefficient0.71 | 6 | |
| Concentration of target information | GPT-2 Small 5 (test) | Gini Coefficient0.72 | 6 | |
| Concentration of target information | GPT-2 Small (test 4) | Gini Coefficient0.74 | 6 | |
| Concentration of target information | GPT-2 Small 3 (test) | Gini Coefficient0.73 | 6 | |
| Concentration of target information | GPT-2 Small (test) | Gini Coefficient0.71 | 6 | |
| Concentration of target information | GPT-2 Small (test 1) | Gini Coefficient0.33 | 6 | |
| Text Generation | GPT-2 Large | Perplexity25.876 | 5 | |
| Machine-generated text detection | GPT-2 (test) | Accuracy85.75 | 5 |