License & Deployment Mix: 17 tools – 9 OSS, 13 SaaS. (OSS and SaaS counts can overlap when an open-source tool also offers a vendor-hosted edition.)
What Is Document Intelligence?
Document intelligence is the layer between a raw scanned page (PDF, image, scan, photo of a form) and the structured data downstream systems need: searchable text, machine-readable tables, key-value pairs, and natural-reading-order markdown suitable for ingestion by an LLM or RAG pipeline.
The category fuses three previously-separate disciplines:
- OCR (optical character recognition) – convert pixels to characters; the classic problem solved by Tesseract since the 1980s
- Layout analysis – detect headers, footers, columns, tables, figures, captions, and reading order so the extracted text is structured, not a wall of noise
- Document understanding – recognize forms, invoices, receipts, equations, handwriting, and produce key-value or schema-bound outputs
Modern document-intelligence tools combine all three using vision-language models (VLMs) and specialised layout / OCR / table-recognition sub-models. The output is typically Markdown, HTML, or JSON suitable for an LLM.
The information on these pages was researched by a combination of human review and large language models. To suggest an addition or correction, please contact us. Prepared by Rhodium Systems Inc., author of the ResorsIT platform — a unified IT operations management platform for IT teams and MSPs that integrates a curated suite of open-source, commercial, and SaaS applications into a single system with shared identity, single sign-on, access control, and a common audit trail. Use this catalogue only as a starting point for your own research, and review any tool carefully against your own requirements before relying on it. Catalogue data version 2026.197.
Comparison
This page compares the 17 document-intelligence tools evaluated in this category across the dimensions that matter most when choosing one.
Quick Reference
| Tool | Origin | Code Licence | Model Licence | Self-Host? | Hosted? |
|---|---|---|---|---|---|
| Marker | Datalab | GPL-3.0 | OpenRAIL-M (revenue cap) | Yes | Yes |
| Surya | Datalab | GPL-3.0 | OpenRAIL-M (revenue cap) | Yes | (via Marker) |
| Chandra | Datalab | Apache 2.0 | OpenRAIL-M (revenue cap) | Yes | Yes |
| Docling | IBM / LF AI | MIT | Per model | Yes | (community) |
| olmOCR | Allen AI | Apache 2.0 | Apache 2.0 | Yes | Demo only |
| Tesseract | Community / Google | Apache 2.0 | n/a | Yes | No |
| PaddleOCR | Baidu | Apache 2.0 | n/a | Yes | No |
| Unstructured | Unstructured-IO | Apache 2.0 (CE) | n/a | Yes (CE) | Yes (Platform) |
| EasyOCR | JaidedAI | Apache 2.0 | n/a | Yes | No |
| LlamaParse | LlamaIndex | Proprietary | – | Enterprise only | Yes (default) |
| Mathpix | Mathpix Inc. | Proprietary | – | Enterprise on-prem | Yes |
| Azure AI Doc Intel | Microsoft | Proprietary | – | Containers | Yes |
| AWS Textract | Amazon | Proprietary | – | No | Yes |
| Google Document AI | Proprietary | – | No | Yes | |
| Mistral OCR | Mistral AI | Proprietary | – | Enterprise | Yes |
| ABBYY | ABBYY | Proprietary | – | Yes (mature) | Cloud SDK |
| Nanonets | Nanonets | Proprietary | – | Enterprise | Yes |
Capability Matrix
| Tool | OCR | Layout | Tables | Equations | Handwriting | Forms | Reading order |
|---|---|---|---|---|---|---|---|
| Marker | Y | Y | Y | Y | Y | Partial | Y |
| Surya | Y | Y | Y | Y | Y | – | Y |
| Chandra | Y | Y | Y | Y | Y (best OSS) | Y | Y |
| Docling | Y | Y | Y | Y | Partial | Partial | Y |
| olmOCR | Y | Y | Y | Y | Y | – | Y |
| Tesseract | Y | – | – | – | Weak | – | Weak |
| PaddleOCR | Y | Y | Y | Y | Partial | Y | Y |
| Unstructured | (delegates) | Y | Y | – | (depends) | – | Y |
| EasyOCR | Y | – | – | – | – | – | – |
| LlamaParse | Y | Y | Y | Y | Y | Y | Y |
| Mathpix | Y | Y | Y | Y (best) | Y (best) | – | Y |
| Azure AI DI | Y | Y | Y | Y | Y | Y | Y |
| AWS Textract | Y | Y | Y | – | Limited | Y | Y |
| Google Doc AI | Y | Y | Y | – | Y | Y | Y |
| Mistral OCR | Y | Y | Y | Y | Y | – | Y |
| ABBYY | Y | Y | Y | – | Y | Y (template) | Y |
| Nanonets | Y | Y | Y | – | Limited | Y | Y |
Deployment Model
| Tool | Local CPU | Local GPU | Docker | SaaS | On-prem enterprise | Air-gap |
|---|---|---|---|---|---|---|
| Marker | Slow | Fast | Yes | Yes | Self-host | Yes |
| Surya | Slow | Fast | Self-build | – | Self-host | Yes |
| Chandra | Very slow | Fast | Yes | Yes | Self-host | Yes |
| Docling | Yes | Faster | Yes | – | Self-host | Yes |
| olmOCR | – | Required | Yes | Demo | Self-host | Yes |
| Tesseract | Yes | – | Yes | – | Self-host | Yes |
| PaddleOCR | Yes | Faster | Yes | – | Self-host | Yes |
| Unstructured | Yes | Yes | Yes | Yes (Platform) | Self-host | Yes |
| EasyOCR | Yes | Yes | Yes | – | Self-host | Yes |
| LlamaParse | – | – | – | Yes | Enterprise | – |
| Mathpix | – | – | – | Yes | Enterprise | Enterprise |
| Azure AI DI | – | – | Disconnected containers | Yes | Yes (subset) | Limited |
| AWS Textract | – | – | – | Yes | – | No |
| Google Doc AI | – | – | – | Yes | – | No |
| Mistral OCR | – | – | – | Yes | Enterprise | Enterprise |
| ABBYY | – | – | – | Cloud SDK | Yes | Yes |
| Nanonets | – | – | – | Yes | Enterprise | Enterprise |
Pricing Model
| Tool | Open source? | Free tier | Paid model |
|---|---|---|---|
| Marker | Yes | Free for users under $2M revenue | Datalab managed: $5 free, then per-page |
| Surya | Yes | Same as Marker | (no separate hosted) |
| Chandra | Yes | Same as Marker | Datalab managed: per-page |
| Docling | Yes | Free | – |
| olmOCR | Yes | Free | – |
| Tesseract | Yes | Free | – |
| PaddleOCR | Yes | Free | – |
| Unstructured | CE only | Free CE | Platform: usage-based |
| EasyOCR | Yes | Free | – |
| LlamaParse | No | 10K free credits | Per-page |
| Mathpix | No | Limited free Snips | Subscription + per-call API |
| Azure AI DI | No | Free tier (low volume) | Per-page tiered |
| AWS Textract | No | (none) | Per-page tiered (3 endpoints) |
| Google Doc AI | No | Free trial | Per-page tiered (per processor) |
| Mistral OCR | No | – | Per-call |
| ABBYY | No | Free trials | Subscription + enterprise per-page |
| Nanonets | No | Free tier | Tiered SaaS + enterprise |
Recommended Defaults
Self-Hosted RAG / Knowledge Base
- Docling – MIT, IBM-backed, no licence anxiety; best default
- Marker – highest practical accuracy / throughput trade-off if revenue under $2M and GPL is acceptable
- olmOCR – when commercial licence restrictions on Marker become a problem
- PaddleOCR – when CPU-only or non-Latin-script (Chinese / Japanese / Korean) accuracy matters
Hosted SaaS for RAG
- LlamaParse – if already using LlamaIndex
- Datalab managed – highest accuracy (Chandra), good batch support
- Mistral OCR – if EU sovereignty matters
- Azure AI / AWS Textract / Google Doc AI – match to whichever cloud the customer already uses
Specialty
- Mathpix – STEM, equations, math- heavy content
- Hyperscience / Rossum / Nanonets / Veryfi – templated business documents (invoices, claims, receipts)
- ABBYY – regulated, on-prem, template-heavy enterprise workflows
Just OCR (no layout)
- Tesseract – portable, ubiquitous, CPU-friendly; the lowest-friction baseline
- EasyOCR – simplest Python API
- PaddleOCR – best accuracy for the trade-off; 80+ languages
Tools
17 tools.
ABBYY FineReader / FlexiCapture / Vantage
ABBYY is the veteran of enterprise OCR. Founded in 1989, the company has been selling commercial OCR for over three decades and remains the default choice for enterprise document-processing workflows where on-prem deployment, regulatory…
License: Proprietary (proprietary) · Kind: web · Deploy: saas, native · SSO: none
AWS Textract
AWS Textract is Amazon’s managed document intelligence service. It covers the full category breadth: text detection (OCR), forms / tables / signatures detection, natural-language queries against documents, and specialist endpoints for **exp…
License: Proprietary (proprietary) · Kind: web · Deploy: saas · SSO: OIDC, SAML
Azure AI Document Intelligence
Azure AI Document Intelligence (formerly Azure Form Recognizer) is Microsoft’s managed document-processing service.
License: Proprietary (proprietary) · Kind: web · Deploy: saas · SSO: OIDC, SAML
Chandra
Chandra is Datalab’s highest-accuracy document intelligence engine – a vision- language model trained specifically for document understanding.
License: Apache-2.0 (OSS) · Kind: web · Deploy: native, docker · SSO: none
Docling
Docling is the document intelligence project from IBM Research Zurich, hosted as a Linux Foundation AI & Data project.
License: MIT (OSS) · Kind: web · Deploy: native · SSO: none
EasyOCR
EasyOCR is the lowest-friction OCR Python library – pip install, three lines of code, and you have OCR working on 80+ languages.
License: Apache-2.0 (OSS) · Kind: web · Deploy: native · SSO: none
Google Document AI
Google Document AI is Google Cloud’s managed document-intelligence service. It exposes processors – pre-built, processor- specific endpoints for a given document type (invoice processor, receipt processor, contract processor, US W-2, U…
License: Proprietary (proprietary) · Kind: web · Deploy: saas · SSO: OIDC, SAML
LlamaParse
LlamaParse is LlamaIndex’s commercial SaaS document-parsing service – the cloud-hosted counterpart to the open-source LlamaIndex RAG framework.
License: Proprietary (proprietary) · Kind: web · Deploy: saas · SSO: OIDC, SAML
Marker
Marker is the headline open-source document intelligence tool from Datalab. It converts PDFs, images, PPTX, DOCX, XLSX, HTML, and EPUB files into structured Markdown, JSON, HTML, or chunked output suitable for RAG ingestion.
License: GPL-3.0-only (OSS) · Kind: web · Deploy: native · SSO: none
Mathpix
Mathpix is the gold standard for STEM and equation OCR. Originally built around the Snip desktop / mobile app – “screenshot an equation, paste LaTeX into your editor” – Mathpix has expanded into a full document-intelligence servic…
License: Proprietary (proprietary) · Kind: web · Deploy: saas · SSO: OIDC, SAML
Mistral OCR API
Mistral OCR is an AI-first document-intelligence API from Mistral AI, converting PDFs and images to Markdown or JSON with preserved layout, tables, and equations and strong multilingual support.
License: Proprietary (proprietary) · Kind: web · Deploy: saas · SSO: OIDC, SAML
Nanonets
Nanonets is a mid-market IDP platform that competes with the hyperscalers’ document-AI services on a friendlier UX and lower price point.
License: Proprietary (proprietary) · Kind: web · Deploy: saas · SSO: OIDC, SAML
olmOCR
olmOCR is the Allen Institute for AI’s contribution to the open-source document intelligence space – a 7B-parameter vision- language model trained specifically to convert PDFs, PNGs, and JPEGs into clean, readable Markdown.
License: Apache-2.0 (OSS) · Kind: web · Deploy: native · SSO: none
PaddleOCR
PaddleOCR is Baidu’s open-source OCR toolkit, built on the PaddlePaddle deep- learning framework.
License: Apache-2.0 (OSS) · Kind: web · Deploy: native · SSO: none
Surya
Surya is Datalab’s OCR + document analysis toolkit – the lower-level building blocks that Marker uses to construct its end-to-end pipeline.
License: GPL-3.0-only (OSS) · Kind: web · Deploy: native · SSO: none
Tesseract
Tesseract is the canonical open-source OCR engine. Originally developed at HP Labs in the 1980s, open-sourced in 2005, sponsored by Google through the 2010s, and now community- maintained, Tesseract is the OCR everyone benchmarks agains…
License: Apache-2.0 (OSS) · Kind: web · Deploy: native · SSO: none
Unstructured
Unstructured.io positions itself explicitly as “ETL for unstructured data” – the piece that lives between a corporate document repository and a vector store / LLM.
License: Apache-2.0 (OSS) · Kind: web · Deploy: native · SSO: none