AI TechnologyAug 10, 2026 23:25 UTC

Legacy OCR Misrecognition Contaminates AI Training Data

Hugging Face and EleutherAI jointly tested 14 open-source OCR models on over 2,000 pages of historical books through the "FineBooks" project. The highest-accuracy model, "dots.mocr", achieved 97.6% character accuracy at a cost below $2 per 1,000 pages. While sufficient for AI training data, it is deemed inadequate for academic transcription.

Legacy OCR Misrecognition Contaminates AI Training Data

Hugging Face and EleutherAI's jointly-led "FineBooks" project conducted a large-scale evaluation of OCR (optical character recognition) technology precision used in digitalizing historical books. The project tested 14 open-source OCR models on scan images of historical texts spanning over 2,000 pages, comparing the performance and practicality of each model.

OCR is a technology that reads characters from scanned images of paper books or documents and converts them into text data. Historical books with old typefaces and frequent printing smudges or distortions are harder to read accurately than modern documents and are prone to errors. When text containing such errors is used as training data for AI, models may learn incorrect language patterns, making data quality assurance a critical challenge.

Test results showed that "dots.mocr" was the most accurate model, achieving 97.6% character-level accuracy. Cost-wise, it was kept below $2 per 1,000 pages, suggesting realistic operational feasibility for large-scale book digitization. However, the project team evaluates this precision as "sufficient for AI training data but insufficient for academic transcription." This indicates that additional refinement is still needed for scenarios requiring word-for-word accuracy, such as in historical research.

The project garners attention against a backdrop of rising awareness regarding data quality in recent AI development. The performance of large language models (LLMs) is influenced not only by the quantity of training data but also by its quality; therefore, text containing typos or character encoding errors can reduce learning effectiveness. Modern text from the internet is already utilized in large quantities, and book and document archives are positioned as promising next data sources.

Digitization of historical books has been an area libraries and research institutions have pursued for years, but quality standards were set for academic use, creating significant cost constraints. The FineBooks approach can be seen as recalibrating the balance between cost and accuracy to realistic levels by focusing specifically on AI training applications. This notion—"adequate for AI purposes even if below academic standards"—could serve as a practical guideline for advancing large-scale data preparation initiatives.

The key question going forward is how well the 97.6% character accuracy will perform in actual AI training environments. Since the impact of the remaining ~2.4% error accumulation across long texts varies by context, continuous validation of effects on actual learning is necessary. Additionally, how historical books' writing styles and vocabularies contribute to linguistic diversity in AI models remains an important consideration.

#OCR#TrainingData#LLM#DataQuality#HuggingFace#OpenSource#DigitalArchive
AI issue Staff

This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.

Comments

Log in to comment