Tag
13 articles
Explore how deepDoctection enables the construction of end-to-end document intelligence pipelines by integrating layout analysis, OCR, and table extraction, with support for custom NER services and RAG workflows.
Learn how to install and use UPDF, a lightweight PDF alternative with AI-powered editing and conversion features.
Learn to build a complete document intelligence pipeline using docTR that performs OCR, layout analysis, KIE, and creates searchable PDFs.
The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on more than 2,000 historical book pages, with the top model achieving 97.6% character accuracy. While this is sufficient for AI training data, it falls short of scholarly transcription standards.
Learn how Datalab Marker v2 improves document OCR speed and accuracy compared to other tools like MinerU, Docling, and LiteParse.
A tutorial from MarkTechPost demonstrates how to build an end-to-end OCR pipeline using Baidu's Unlimited-OCR, enabling high-resolution image and multi-page PDF processing with advanced features like tiled inference and cross-page content handling.
Learn how Baidu's Unlimited OCR achieves efficient processing of dozens of document pages in a single pass by mimicking human memory and forgetting mechanisms.
Learn how OCRmyPDF converts scanned documents into searchable PDFs using Optical Character Recognition (OCR) technology, and how it helps organize and access information more efficiently.
Mistral AI's new OCR 4 model outperforms competitors in 72 percent of blind test cases, according to the company. The model is designed to extract text from various document formats including PDFs, Word files, and PowerPoint presentations.
Baidu open-sources Unlimited OCR, a 3B-parameter model that maintains flat KV cache for efficient long-document parsing, scoring 93.23 on OmniDocBench v1.5.
Mistral AI's Mistral OCR 4 introduces citation-ready, structured document outputs that enhance RAG, agentic, and enterprise search pipelines. The model supports 170 languages and runs in a single self-hosted container.
Baidu's Qianfan-OCR is a 4B-parameter unified document intelligence model that streamlines document processing by combining layout analysis, parsing, and understanding into a single vision-language architecture.