Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Decoupling Semantics from Distortions: Multi-Scale Two-Stream Vision-Language Alignment for AI-Generated Image Quality Assessment

About

Existing vision-language model (VLM)-based AI-generated image quality assessment (AIGIQA) methods suffer from a fundamental semantic-distortion dimensional conflict: monolithic representations optimized for semantic discrimination inherently entangle compositional understanding with low-level perceptual sensitivity, rendering them blind to fine-grained quality degradations. We introduce MST-CLIPIQA, a multi-scale two-stream framework that achieves hierarchical vision-language alignment through explicit representational decoupling. Our architecture leverages dual CLIP encoders with complementary patch granularities: coarse-grained streams capture global semantic coherence while fine-grained streams preserve textural signatures and artifact patterns. An information bottleneck-inspired gated fusion mechanism performs adaptive cross-scale distillation, with optional cross-attention enabling prompt-anchored correspondence evaluation when generation prompts are available. Extensive experiments across five benchmarks establish new state-of-the-art results, achieving average improvements of 1.11 percent SRCC on quality and 2.35 percent SRCC on text-image correspondence prediction, while maintaining efficiency with only 0.8M trainable parameters. Our project is available at https://github.com/YMlinfeng/MST-CLIPIQA.

Zijie Meng• 2026

Related benchmarks

TaskDatasetResultRank
Image Quality AssessmentAGIQA-3K
SRCC0.9091
175
Image Quality AssessmentAGIQA-1K
SRCC0.9091
68
Visual Quality AssessmentAIGIQA-20K
SRCC0.8936
31
Authenticity score predictionAIGCIQA 2023
SRCC87.01
25
Authenticity score predictionPKU-AIGIQA-4K
SRCC0.8289
22
Showing 5 of 5 rows

Other info

Follow for update