| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Image Captioning | TextCaps | CIDEr164.3 | 154 | |
| Image Captioning | TextCaps (test) | CIDEr164.3 | 82 | |
| Image Captioning | TextCaps (val) | CIDEr163.7 | 79 | |
| Image Captioning | TextCaps K = 3 (test) | mBLEU-478.6 | 12 | |
| Text-oriented Visual Question Answering | TextCaps | CIDEr144.9 | 7 | |
| Image Reconstruction | TextCaps (test) | FID15.51 | 6 | |
| Image-to-Text Retrieval | TextCaps | Recall@197.4 | 4 | |
| Text-to-Image Retrieval | TextCaps | Recall@189.6 | 4 | |
| Visually Grounded Language Generation | TextCaps (test) | Score152 | 4 |