Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

Typhoon OCR: Open Vision-Language Model For Thai Document Extraction

About

Document extraction is a core component of digital workflows, yet existing vision-language models (VLMs) predominantly favor high-resource languages. Thai presents additional challenges due to script complexity from non-latin letters, the absence of explicit word boundaries, and the prevalence of highly unstructured real-world documents, limiting the effectiveness of current open-source models. This paper presents Typhoon OCR, an open VLM for document extraction tailored for Thai and English. The model is fine-tuned from vision-language backbones using a Thai-focused training dataset. The dataset is developed using a multi-stage data construction pipeline that combines traditional OCR, VLM-based restructuring, and curated synthetic data. Typhoon OCR is a unified framework capable of text transcription, layout reconstruction, and document-level structural consistency. The latest iteration of our model, Typhoon OCR V1.5, is a compact and inference-efficient model designed to reduce reliance on metadata and simplify deployment. Comprehensive evaluations across diverse Thai document categories, including financial reports, government forms, books, infographics, and handwritten documents, show that Typhoon OCR achieves performance comparable to or exceeding larger frontier proprietary models, despite substantially lower computational cost. The results demonstrate that open vision-language OCR models can achieve accurate text extraction and layout reconstruction for Thai documents, reaching performance comparable to proprietary systems while remaining lightweight and deployable.

Surapon Nonesung, Natapong Nitarach, Teetouch Jaknamon, Pittawat Taveekitworachai, Kunat Pipatanakul• 2026

Related benchmarks

TaskDatasetResultRank
Thai document parsingIn-house Thai document corpus Government Forms
Levenshtein Error Rate0.035
10
Thai document parsingIn-house Thai document corpus Financial Reports
BLEU0.91
6
Thai document parsingIn-house Thai document corpus Thai Books
BLEU64
6
Document TranscriptionThai Document Understanding Evaluation
ROUGE-L (Thai Books)94.9
4
Document TranscriptionThai Documents Books
Levenshtein Distance0.053
4
Document TranscriptionThai Documents Thai Financial Reports
Levenshtein Distance0.079
4
Document TranscriptionThai Documents Average
Levenshtein Distance0.251
4
Optical Character RecognitionThai Books Typhoon OCR (test)
BLEU0.746
4
Optical Character RecognitionThai Government Forms Typhoon OCR (test)
BLEU Score0.87
4
Optical Character RecognitionThai Financial Reports Typhoon OCR (test)
BLEU Score0.849
4
Showing 10 of 17 rows

Other info

GitHub

Follow for update