Challenge: Existing methods for historical document restoration focus on single modality or limited-size restoration, failing to meet practical needs.
Approach: They propose a full-page HDR dataset and an automated HDR solution to replace manual restoration methods.
Outcome: The proposed solution improves OCR accuracy from 46.83% to 84.05% when processing severely damaged documents, with enhancement to 94.25% through human-machine collaboration.

Similar Papers

Draft, Verify, Restore: Self-Refining Historical Inscription Restoration with a Unified MLLM (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for end-to-end historical inscription restoration rely on task-separated pipelines with irreversible error accumulation and patch-based generation that sacrifices page-level consistency.
Approach: They propose a unified MLLM for end-to-end historical inscription restoration that integrates draft-guided localization and Hierarchical self-refinement to enable accurate damage localization.
Outcome: The proposed model achieves superior performance in both text restoration accuracy and appearance restoration quality.
Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Historical documents suffer from illegibility due to physical deterioration and damage due to deteriorating materials.
Approach: a new framework leverages large language models with retrieval-augmented generation to restore historical documents. authors propose a framework that leverages implicit knowledge of pre-trained LLMs with explicitly retrieved external context.
Outcome: a new framework outperforms existing methods for restoration of historical documents in Korean . the proposed model can restore both general characters and named entities, the authors say .
PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy (2025.acl-long)

Copied to clipboard

Challenge: Existing pipelines that combine document image restoration with semantic-aware post-OCR correction can improve text extraction from degraded images.
Approach: They propose a two-stage pipeline that combines document image restoration with semantic-aware post-OCR correction to enhance both visual clarity and textual consistency.
Outcome: The proposed pipeline reduces character error rates by 63.9-70.3% on 13,831 pages of real historical documents in English, French, and Spanish compared to OCR on raw images.
HisDoc-OCR: Restoring Visual Grounding in MLLMs for Chinese Historical Document OCR (2026.findings-acl)

Copied to clipboard

Challenge: Despite multimodal large language models' strong performance on modern document OCR, their application to historical Chinese texts suffers from severe hallucinations, character fabrication, uncontrolled repetition, and semantic drift.
Approach: They propose a multimodal large language model which restores visual grounding through three synergistic strategies: Layout Injection, First-Occurrence Boost, Self-Distilled Attention Focusing and HisDoc-OCR.
Outcome: The proposed model outperforms general-purpose and OCR-specific models on Chinese historical documents.
Restoring and Mining the Records of the Joseon Dynasty via Neural Language Modeling and Machine Translation (2021.naacl-main)

Copied to clipboard

Challenge: voluminous historical records are difficult to fully utilize since they are written in ancient languages and some parts are damaged over time.
Approach: They propose a multi-task learning approach to restore and translate historical documents using a self-attention mechanism.
Outcome: The proposed approach improves the accuracy of the translation task over baselines without multi-task learning.
Efficient OCR for Building a Diverse Digital History (2024.acl-long)

Copied to clipboard

Challenge: Current optical character recognition (OCR) systems are poorly extensible to low-resource document collections, as learning a language-vision model requires extensive labeled sequences and compute.
Approach: They propose to model optical character recognition as a character level image retrieval problem using a contrastively trained vision encoder.
Outcome: The proposed model is more sample efficient and extensible than existing architectures, enabling accurate OCR in settings where existing solutions fail.
PHD: Pixel-Based Language Modeling of Historical Documents (2023.emnlp-main)

Copied to clipboard

Challenge: Recent years have seen a boom in efforts to digitise historical documents in numerous languages and sources, leading to a transformation in the way historians work.
Approach: They propose a method for generating synthetic scans to resemble real historical documents by pre-training a model to reconstruct masked patches instead of predicting token distributions.
Outcome: The proposed model can reconstruct masked patches and understand language well.
EfficientOCR: An Extensible, Open-Source Package for Efficiently Digitizing World Knowledge (2023.emnlp-demo)

Copied to clipboard

Challenge: Existing OCR engines fail to provide accurate, cost-effective and sample-efficient character recognition for public domain documents.
Approach: EffOCR is an open-source optical character recognition package that is accurate, cheap to deploy and sample efficient to customize to novel collections, languages, and character sets.
Outcome: EffOCR model trains character retrieval problem and scales to novel collections, languages, and character sets.
CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition (2025.emnlp-main)

Copied to clipboard

Challenge: Existing vision-language models are not equipped to read diverse languages and scripts found in historical materials.
Approach: They propose to train an open-weight vision-language model for historical text recognition on CHURRO-DS, the largest historical text-recognition dataset to date.
Outcome: The proposed model outperforms existing vision-language models on CHURRO-DS, the largest historical text recognition dataset to date.
A Large-Scale Comparison of Historical Text Normalization Systems (N19-1)

Copied to clipboard

Challenge: a large study of historical text normalization is done on eight languages . there is no consensus on the state-of-the-art approach to normalization .
Approach: They present a large study of historical text normalization done on eight languages . they evaluate four different systems based on supervised learning on datasets from eight different languages based in the literature .
Outcome: The proposed methods are based on supervised learning and are available online.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations