Research & Papers

ICDAR 2026 HIPE-OCRepair shows LLMs fix historical OCR errors with caveats

LLMs boost OCR accuracy on 17th–20th century texts but risk over-correcting clean pages

Deep Dive

Researchers from academia and industry—including Maud Ehrmann, Emanuela Boros, Juri Opitz, and Simon Clematide—organized the ICDAR 2026 HIPE-OCRepair competition to benchmark LLM-assisted OCR post-correction on historical documents. The task required correcting noisy OCR transcripts from digitized historical newspapers and printed works spanning the 17th to 20th centuries across three languages (English, French, German). Participants worked at the level of coherent transcription units (paragraphs or articles) without access to source images, simulating real-world digital heritage workflows.

Four teams submitted approaches ranging from zero-shot prompting of off-the-shelf LLMs to continued pre-training and fine-tuning on domain-specific data. The evaluation adopted a retrieval-oriented scoring method, focusing on search and access utility rather than exact diplomatic transcription. Results demonstrated that modern LLMs can significantly improve OCR quality, but performance varied widely across datasets, languages, and noise conditions. A recurring issue was over-correction on already clean or low-noise inputs, where the LLM introduced hallucinations. The competition released a harmonized multilingual dataset, scorer, and evaluation pipeline to support future research.

Key Points
  • Four teams competed with strategies from zero-shot prompting to fine-tuning LLMs for historical text correction
  • Dataset covers English, French, and German newspapers and books from the 17th–20th century with varied noise levels
  • Over-correction on low-noise inputs emerged as a major challenge, highlighting limits of LLM-based post-correction

Why It Matters

LLMs can accelerate digitizing historical archives, but careful tuning is needed to avoid introducing errors.

📬 Get the top 10 AI stories daily