Docling ranks #1 of 9 in our document extraction apis testing. free forever, permissive on both the code and the models, and nothing ever leaves your infrastructure..
90/100
the only entry where the privacy question is answered by architecture rather than by a policy page — and it publishes real accuracy figures while doing it.
why people look for an alternative
−no vendor sla, support or managed scaling
−two separate licences to accept, code and models
−accuracy and throughput depend on your own hardware
−no hosted option if you don't want to run it
stay with Docling if free with no tiers, mit code and permissive model licences is the thing you care about most — nothing below beats it on that.
#2 in document extraction apis · zero data retention on every plan including the free one, and a bill that stops at $3,000 however much you send.
87/100
verdictthe strongest default data posture of any hosted option here, on pricing with a genuinely unusual ceiling — expensive per page until you're sending a lot, then free.
Unstructured vs Docling
Docling
Unstructured
price
free
$30 / 1,000 pages
free tier
yes
yes
cost/1000 pages
$0
$30
tables cost extra
no
no
licence
mit code, cdla-permissive models
apache 2.0 core
trains on your docs
no — self-hosted
no — zero retention
self-host
yes, only option
yes, open core
switch forteams handling sensitive documents at volume who want a hosted service without a retention conversation.
pros
+zero data retention stated on all plans, including free
+15,000 free pages a month, no card required
+spend caps at $3,000/month, then free to a million pages
+apache 2.0 core library, self-hostable
cons
−$30 per thousand pages is high at low volume
−flat rate gives no visibility into how hard documents are handled
#5 in document extraction apis · the best table extraction here, proven on a benchmark it built and published rather than one it cites.
78/100
verdictthe strongest table story in the category and the least predictable bill — the same thousand pages can cost $15 or $60 and you find out afterwards.
Reducto vs Docling
Docling
Reducto
price
free
$15 / 1,000 pages
free tier
yes
yes
cost/1000 pages
$0
$15-60 by complexity
tables cost extra
no
pushes to higher credit band
licence
mit code, cdla-permissive models
proprietary api
trains on your docs
no — self-hosted
not confirmed
self-host
yes, only option
enterprise, unverified
switch fordense financial and scientific documents where table structure is the whole job.
pros
+publishes rd-tablebench, an open hand-labelled dataset, with its score
+purpose-built for complex tables and scans
+15,000 free credits a month
+trust centre references on-prem and air-gapped options
cons
−cost per thousand pages swings fourfold by content
#8 in document extraction apis · the cheapest plain ocr here, the most expensive forms extraction, and it trains on your documents unless you stop it.
66/100
verdictthe most granular published pricing in the category, and the only vendor here whose default is to use your contracts and invoices to improve its own models.
AWS Textract vs Docling
Docling
AWS Textract
price
free
$1.50 / 1,000 pages
free tier
yes
yes
cost/1000 pages
$0
$1.50 text, $50 forms
tables cost extra
no
yes — 10x base rate
licence
mit code, cdla-permissive models
proprietary aws service
trains on your docs
no — self-hosted
yes, unless you opt out
self-host
yes, only option
no
switch foraws-native pipelines doing plain text extraction at volume, with the training opt-out configured first.
pros
+cheapest plain ocr at $1.50 per thousand pages
+every feature separately and clearly priced
+volume discounts above a million pages a month
+specialised models for expense, id and lending
cons
−documents used to improve aws models unless you opt out
−opt-out is an organizations-level policy, not a setting
every tool on this page went through the same test as Docling — same tasks, same order, scored the same way. the comparison tables are the figures from that testing, not vendor spec sheets.