olmOCR 2: Unit Test Rewards for Document OCR

About

We present olmOCR 2, the latest in our family of powerful OCR systems for converting digitized print documents, like PDFs, into clean, naturally ordered plain text. olmOCR 2 is powered by olmOCR-2-7B-1025, a specialized, 7B vision language model (VLM) trained using reinforcement learning with verifiable rewards (RLVR), where our rewards are a diverse set of binary unit tests. To scale unit test creation, we develop a pipeline for generating synthetic documents with diverse and challenging layouts, known ground-truth HTML source code, and extracted test cases. We show that RL training on these test cases results in state-of-the-art performance on olmOCR-Bench, our English-language OCR benchmark, with the largest improvements in math formula conversion, table parsing, and multi-column layouts compared to previous versions. We release our model, data and code under permissive open licenses.

Jake Poznanski, Luca Soldaini, Kyle Lo• 2025

Related benchmarks

Task	Dataset	Result
Document Parsing	olmOCR-bench	ArXiv Processing Accuracy82.9	59
Document Parsing	OmniDocBench	Overall Accuracy80.03	32
PDF-to-LaTeX reconstruction	TEXOCR-Bench 1.0 (test)	Section Accuracy (SA)43.7	23
Table Extraction	100 pages (451 tables) synthetic (test)	LLM Score (Overall)4.05	21
Full-page OCR	English Fox	Page CER1.8	12
Color-guided OCR	English Fox	Color CER55.8	12
Line-level OCR	English Fox	Line CER87.8	12
Region-level OCR	English Fox	Region CER69.7	12
Grounded OCR	OCR-IDL, TabMe++, and PubMed-OCR (10.5K held-out pages)	CER (text)0.365	11

Showing 9 of 9 rows

Other info

Follow for update

@wizwand_team Discord