Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

TikZero: Zero-Shot Text-Guided Graphics Program Synthesis

About

Automatically synthesizing figures from text captions is a compelling capability. However, achieving high geometric precision and editability requires representing figures as graphics programs in languages like TikZ, and aligned training data (i.e., graphics programs with captions) remains scarce. Meanwhile, large amounts of unaligned graphics programs and captioned raster images are more readily available. We reconcile these disparate data sources by presenting TikZero, which decouples graphics program generation from text understanding by using image representations as an intermediary bridge. It enables independent training on graphics programs and captioned images and allows for zero-shot text-guided graphics program synthesis during inference. We show that our method substantially outperforms baselines that can only operate with caption-aligned graphics programs. Furthermore, when leveraging caption-aligned graphics programs as a complementary training signal, TikZero matches or exceeds the performance of much larger models, including commercial systems like GPT-4o. Our code, datasets, and select models are publicly available.

Jonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka, Masao Utiyama, Raj Dabre, Steffen Eger, Simone Paolo Ponzetto• 2025

Related benchmarks

TaskDatasetResultRank
TikZ code generationDaTikZ V4 (evaluation)
CLIP Similarity0.104
15
Scientific Figure GenerationFigureBench Paper category
Aesthetic Quality2
13
Showing 2 of 2 rows

Other info

Follow for update