Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Reading Order Inference for Complex Document Layouts

About

Reading order inference remains a critical bottleneck in the digitization of complex historical manuscripts, where pages contain multiple spatially interleaved reading streams, the canonical example being the Glossa Ordinaria layout, in which a central text is surrounded by commentaries that wrap around it in non-rectangular, non-convex regions. We present a training-free, graph-based framework: each OCR text line becomes a node in a directed candidate-transition graph, edges are scored by a weighted additive ensemble of two lightweight language-model signals (causal language model conditional likelihood and BERT next-sentence prediction, NSP; a third sentence-embedding signal was evaluated but did not improve reading order), and the global reading order is recovered as a degree-constrained directed path cover. To avoid the cascading "edge-theft" failures of greedy edge selection, we propose a max-regret inference rule that prioritizes commitments with high opportunity cost. We evaluate on synthetic Glossa Ordinaria grid layouts, on 23 ALTO page geometries (10 historical source pages plus mirrored and flipped variants), and on a 140-page multi-column English subset of OmniDocBench, comparing our method against the canonical recursive XY-cut (PaddleOCR PP-StructureV3) and two LayoutReader variants (layout-only and text+layout) on identical inputs. On wrap-around Glossa layouts our method recovers 95% of ground-truth successor edges on average vs. XY-cut's 50%; on the OmniDocBench multi-column subset it reaches 88% macro edge accuracy versus XY-cut's 75% and LayoutReader's 25%. The LayoutReader baselines transfer poorly due to a word-level vs. line-level granularity mismatch. We additionally verify mirror-invariance under horizontal and vertical page reflections: Our method changes by less than 1 percentage point, classical XY-cut by 2 points, and LayoutReader-T by up to 8 points.

Iddo Hakim, Sharva Gogawale, Omer Ventura, Gal Grudka, Daria Vasyutinsky-Shapira, Berat Kurar-Barakat, Nachum Dershowitz• 2026

Related benchmarks

TaskDatasetResultRank
Reading Order InferenceALTO 11-page corpus Wrap-around Glossa n=4
Edge Accuracy94.8
4
Reading Order InferenceALTO 11-page corpus (Entire corpus)
Edge Accuracy95.4
4
Reading Order InferenceALTO 11-page corpus n=4 (Manhattan-ceiling)
Edge Accuracy96
4
Reading Order InferenceALTO 11-page corpus Mixed near-Manhattan regime n=3
Edge Accuracy95.3
4
Reading Order PredictionOmniDocBench English multi-column academic_lit
Edge Accuracy93.4
3
Reading Order PredictionOmniDocBench English multi-column (exam paper)
Edge Accuracy96.9
3
Reading Order PredictionOmniDocBench English multi-column magazine
Edge Accuracy89.1
3
Reading Order PredictionOmniDocBench English multi-column (Macro)
Edge Accuracy88
3
Reading Order PredictionOmniDocBench English multi-column (newspaper)
Edge Accuracy74.2
3
Reading Order PredictionOmniDocBench English multi-column (book)
Edge Accuracy82
3
Showing 10 of 11 rows

Other info

Follow for update