| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Consistent Text-to-Image Generation | ConsiStory+ (test) | CLIP-T0.9074 | 23 | |
| Visual Storytelling | ConsiStory+ | CLIP-T Score0.8584 | 7 | |
| Consistent text-to-image generation | ConsiStory+ | CLIP-T0.8889 | 3 | |
| Story Editing | ConsiStory+ protocol | Consistency89 | 3 |