Unstructured ranks #2 of 9 in our document extraction apis testing. zero data retention on every plan including the free one, and a bill that stops at $3,000 however much you send..
87/100
the strongest default data posture of any hosted option here, on pricing with a genuinely unusual ceiling — expensive per page until you're sending a lot, then free.
why people look for an alternative
−$30 per thousand pages is high at low volume
−flat rate gives no visibility into how hard documents are handled
−no published accuracy benchmark
−business tier pricing is custom
stay with Unstructured if zero data retention stated on all plans, including free is the thing you care about most — nothing below beats it on that.
#1 in document extraction apis · free forever, permissive on both the code and the models, and nothing ever leaves your infrastructure.
90/100
verdictthe only entry where the privacy question is answered by architecture rather than by a policy page — and it publishes real accuracy figures while doing it.
Docling vs Unstructured
Unstructured
Docling
price
$30 / 1,000 pages
free
free tier
yes
yes
cost/1000 pages
$30
$0
tables cost extra
no
no
licence
apache 2.0 core
mit code, cdla-permissive models
trains on your docs
no — zero retention
no — self-hosted
self-host
yes, open core
yes, only option
switch foranyone processing documents they cannot send to a third party, which in this category is most people.
pros
+free with no tiers, mit code and permissive model licences
+nothing leaves your infrastructure — no retention question
+published accuracy figures in a technical report
+docling-serve turns it into your own api
cons
−no vendor sla, support or managed scaling
−two separate licences to accept, code and models
−accuracy and throughput depend on your own hardware
#5 in document extraction apis · the best table extraction here, proven on a benchmark it built and published rather than one it cites.
78/100
verdictthe strongest table story in the category and the least predictable bill — the same thousand pages can cost $15 or $60 and you find out afterwards.
Reducto vs Unstructured
Unstructured
Reducto
price
$30 / 1,000 pages
$15 / 1,000 pages
free tier
yes
yes
cost/1000 pages
$30
$15-60 by complexity
tables cost extra
no
pushes to higher credit band
licence
apache 2.0 core
proprietary api
trains on your docs
no — zero retention
not confirmed
self-host
yes, open core
enterprise, unverified
switch fordense financial and scientific documents where table structure is the whole job.
pros
+publishes rd-tablebench, an open hand-labelled dataset, with its score
+purpose-built for complex tables and scans
+15,000 free credits a month
+trust centre references on-prem and air-gapped options
cons
−cost per thousand pages swings fourfold by content
#8 in document extraction apis · the cheapest plain ocr here, the most expensive forms extraction, and it trains on your documents unless you stop it.
66/100
verdictthe most granular published pricing in the category, and the only vendor here whose default is to use your contracts and invoices to improve its own models.
AWS Textract vs Unstructured
Unstructured
AWS Textract
price
$30 / 1,000 pages
$1.50 / 1,000 pages
free tier
yes
yes
cost/1000 pages
$30
$1.50 text, $50 forms
tables cost extra
no
yes — 10x base rate
licence
apache 2.0 core
proprietary aws service
trains on your docs
no — zero retention
yes, unless you opt out
self-host
yes, open core
no
switch foraws-native pipelines doing plain text extraction at volume, with the training opt-out configured first.
pros
+cheapest plain ocr at $1.50 per thousand pages
+every feature separately and clearly priced
+volume discounts above a million pages a month
+specialised models for expense, id and lending
cons
−documents used to improve aws models unless you opt out
−opt-out is an organizations-level policy, not a setting
every tool on this page went through the same test as Unstructured — same tasks, same order, scored the same way. the comparison tables are the figures from that testing, not vendor spec sheets.