Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin

About

Recent advances in Automatic Text Recognition (ATR) have improved access to historical archives, yet a methodological divide persists between palaeographic transcriptions and normalized digital editions. While ATR models trained on more palaeographically-oriented datasets such as CATMuS have shown greater generalizability, their raw outputs remain poorly compatible with most readers and downstream NLP tools, thus creating a usability gap. On the other hand, ATR models trained to produce normalized outputs have been shown to struggle to adapt to new domains and tend to over-normalize and hallucinate. We introduce the task of Pre-Editorial Normalization (PEN), which consists in normalizing graphemic ATR output according to editorial conventions, which has the advantage of keeping an intermediate step with palaeographic fidelity while providing a normalized version for practical usability. We present a new dataset derived from the CoMMA corpus and aligned with digitized Old French and Latin editions using passim. We also produce a manually corrected gold-standard evaluation set. We benchmark this resource using ByT5-based sequence-to-sequence models on normalization and pre-annotation tasks. Our contributions include the formal definition of PEN, a 4.66M-sample silver training corpus, a 1.8k-sample gold evaluation set, and a normalization model achieving a 6.7% CER, substantially outperforming previous models for this task.

Thibault Cl\'erice, Rachel Bawden, Anthony Glaise, Ariane Pinche, David Smith• 2026

Related benchmarks

TaskDatasetResultRank
Pre-Editorial NormalizationGold data manual curated (test)
CER (All)6.7
3
Pre-Editorial NormalizationSchonhardt data 2025 (first 2,000 entries)
CER0.111
3
LemmatizationVie de Saint Lambert BnF Fr. 412 (All tokens)
Precision90.9
2
LemmatizationVie de Saint Lambert 100 most frequent words (MFW) BnF Fr. 412
Precision94.3
2
POS 3-grams TaggingVie de Saint Lambert BnF Fr. 412 (All tokens)
Precision87.7
2
POS TaggingVie de Saint Lambert BnF Fr. 412 (All tokens)
Precision98
2
Showing 6 of 6 rows

Other info

Follow for update