Language models trained on historical texts are facing a significant bottleneck: poor-quality OCR (Optical Character Recognition) data. A new initiative called FineBooks, led by Hugging Face and EleutherAI, is aiming to address this issue by improving the accuracy of scanned historical documents at scale.
Testing OCR Models on Historical Texts
The FineBooks project evaluated 14 open-source OCR models on over 2,000 pages from historical books. The results revealed that while current OCR technology has made significant strides, it still struggles with older fonts, degraded text, and complex layouts—common in historical documents. The top-performing model, dots.mocr, achieved a character accuracy rate of 97.6 percent, at a cost of less than two dollars per thousand pages. This level of accuracy is sufficient for training large language models, but falls short of the precision required for scholarly transcription.
Implications for AI and Historical Research
High-quality text data is crucial for training AI models that can understand and generate human-like language. However, when models are trained on low-quality OCR text, they often inherit inaccuracies, which can skew results and limit the model's ability to perform nuanced tasks like summarization or question answering. The FineBooks team emphasizes that while their approach can significantly enhance data quality for AI training, it's not yet sufficient for academic or archival purposes that demand perfect fidelity.
By focusing on scalable solutions, the project aims to unlock vast repositories of historical texts that are currently underutilized due to OCR limitations. This effort could have wide-ranging impacts, from improving historical AI research to enabling better digital humanities initiatives.
Conclusion
The FineBooks project marks a significant step forward in preparing historical documents for modern AI applications. While there's still work to be done to achieve perfect transcription accuracy, the initiative shows that with the right tools and methodologies, it's possible to dramatically improve the quality of OCR data at scale.



