AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model
About
Large Vision and Language Models (LVLMs) have advanced rapidly, yet European Portuguese (pt-PT) remains systematically underserved by existing open-source multimodal models, which either conflate it with Brazilian Portuguese or severely under-represent it in their training data mixes. We introduce AMALIA-VL, the first open-source instruction-tuned LVLM built natively for pt-PT, pairing a high-resolution vision encoder with dynamic image tiling and a fully open pt-PT-optimized language model via a learned connector. We contribute with a purposefully designed three-stage training process - vision-language alignment, general visual instruction tuning, and preference optimization - together with a pt-PT-centric multimodal data mix combining curated and translated public datasets with novel datasets that address the near-total absence of European Portuguese multimodal resources. Our evaluation shows that AMALIA-VL establishes a strong baseline for open-source pt-PT LVLMs. We will release model weights, training data, and construction pipelines along with machine-translated pt-PT evaluation benchmarks to help democratize pt-PT LVLM development.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Captioning | COCO pt-PT | Accuracy50.8 | 16 | |
| General VQA | POPE pt-PT | Accuracy89.6 | 16 | |
| Captioning | RefCap pt-PT | Accuracy11.5 | 16 | |
| OCR & Document Understanding | TxtVQA pt-PT | Accuracy69.5 | 16 | |
| Spatial Understanding | RefRec pt-PT | Accuracy81.3 | 16 | |
| OCR & Document Understanding | DocVQA pt-PT | Accuracy69.7 | 16 | |
| Chart & Diagram Understanding | ChartQA pt-PT | Accuracy67.7 | 16 | |
| General VQA | SEED pt-PT | Accuracy73.7 | 16 | |
| General VQA | RWQA pt-PT | Accuracy52.2 | 16 | |
| OCR & Document Understanding | OCR pt-PT | Accuracy63.7 | 16 |