Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Aligned but Stereotypical? How System Prompts Shape Demographic Bias in LLM-Based Text-to-Image Models

About

Text-to-image (T2I) systems increasingly rely on Large Language Model (LLM)-based text conditioning to interpret and expand user prompts. While this improves prompt understanding and text-image alignment, we find that it can also introduce implicit demographic assumptions, even when demographic attributes are unspecified. To systematically investigate this behavior across varying levels of prompt ambiguity and complexity, we construct a comprehensive benchmark covering diverse prompt settings. Evaluations on eight recent T2I models show that LLM-based systems consistently exhibit stronger demographic skew than non-LLM-based baselines. We further analyze system prompts, a component unique to LLM-based T2I systems that guides prompt interpretation and expansion. Our analyses show that these instructions strongly influence text embeddings, which subsequently leads to biased image generations. Motivated by these findings, we propose FairPro, a training-free debiasing framework that adaptively generates fairness-aware instructions while preserving user intent. Experiments demonstrate that FairPro substantially reduces demographic disparities while maintaining prompt fidelity.

NaHyeon Park, Na Min An, Kunhee Kim, Soyeon Yoon, Jiahao Huo, Hyunjung Shim• 2025

Related benchmarks

TaskDatasetResultRank
Image Generation DiversityPrompt Complexity (Rewritten)
CLIP Similarity0.9235
6
Image Generation DiversityPrompt Complexity (Simple)
CLIP Score88.39
6
Image Generation DiversityPrompt Complexity (Context)
CLIP Similarity90.38
6
Image Generation DiversityPrompt Complexity Mean
CLIP Similarity0.8919
6
Image Generation DiversityPrompt Complexity (Occupation)
CLIP Score0.8563
6
Bias MeasurementFull Prompt Benchmark Aggregate (test)
Gender Bias Score81.6
6
Social Bias EvaluationTIBET
Gender87
6
Text-to-Image GenerationOccupation prompts
Bias0.746
4
Text-to-Image GenerationSimple prompts
Bias0.797
4
Text-to-Image GenerationContext prompts
Bias0.815
4
Showing 10 of 11 rows

Other info

Follow for update