
Authors

The prediction was that one general model would absorb OCR and IDP. Instead the category exploded, and the reason points at where the actual work is now.
One thing we hear in nearly every sales call is some version of this: "Frontier models keep getting better, won't they just swallow Unstructured whole?"
It’s a fair question, and for a while it looked like the obvious future: drop a PDF into a capable enough model and the whole discipline of OCR and document processing quietly goes away.
It has not worked out that way. The market for intelligent document processing was worth a little over two billion dollars in 2024 and is growing at roughly a third every year. A category that was supposed to be dissolving is somehow one of the faster-growing corners of enterprise software right now.
The work is older than the hype, and the hype just raised its stakes.
Document processing did not begin with GPT-4V. One of its first famous wins was a neural network that read the handwriting on bank checks. Back in the 1990s, Yann LeCun's team at Bell Labs built neural networks that read the handwriting on bank checks. (the models doing it were a few hundred kilobytes then; today's are hundreds of gigabytes 🙂).
The modern lineage picked up around 2020, when LayoutLM learned a document's text and its two-dimensional layout together. This has been a serious field for a long time. What the frontier models changed was not the demand. It was the stakes.
Once RAG and agents arrived, the quality of your extracted data started deciding the quality of everything after it. A garbled table becomes garbled context, and garbled context becomes a confident, wrong answer that nobody catches until it blows up. And this is happening in production now.
Reading the page got easy. Trusting every field did not.
The hard part has kept moving, and you can trace where it went:
- 2020 to 2023 was about reading a document at all. Could a model find the structure, hold the layout, and turn a messy page into something usable.
- 2024 to 2025 was about attention and hallucination. Whether a model could get through a long document without losing the middle or inventing values that were never there.
- 2026 is about precision, where being a little wrong is the same as being wrong. Raw reading is largely handled now. On OmniDocBench, a standard document-parsing benchmark, the strongest frontier models now score around 90 overall, while their table and reading-order scores sit noticeably lower. That is good. It is not good enough when the output feeds an automated decision on an insurance claim.
This is also why the field kept building specialized models instead of standing down. Through 2025 a wave of purpose-built document models appeared next to the general ones. If a general model were enough on its own, that would not be happening.
Why the next model won't simply absorb it.
The models will keep improving, and the bar they have to clear rises right along with them. As extraction gets trustworthy enough to rely on, people put it into higher-stakes loops, agents acting on the data with no human in between, so the precision the job demands climbs faster than raw capability closes the gap. The value keeps concentrating at the precision-critical, verifiable last mile.
None of this is an argument against frontier models. A frontier model is a powerful component, and we use one. On its own it gets you a demo. Trusting it in business-critical work is a bigger ask. It has to read every file type you throw at it, keep up with your volume, and get each field right, because a single wrong value carries all the way downstream. That is the bet we have made at Unstructured.
So what should a document tool do now?
The next time you talk to a document vendor, hold them to three things.
First, every part of a page is its own problem. A table, a form, a chart, handwriting, a footnote tied to a cell: each has its own quirks and needs handling shaped to it. A general model treats the whole page as one job. The tools worth using have put in the time to learn where each part breaks, and built for it.
Second, you should be able to trace the output. Every element should come back with a bounding box, the coordinates of where it sat on the page. When a value matters, you can point to exactly where it came from and check it against the source. (Claude does not give you coordinates for texts it picked up)
Third, it should be a chain you can compose. Some teams want one endpoint: a document in, structured data out. Others need the whole journey, picking data up where it lives and landing it somewhere downstream. Take the slice you need now, and grow into the rest.
Where you should go next!
The prediction was that documents would stop being a problem. Instead they became a more valuable one, and the hard part moved from reading the page to trusting every field on it.
That part is not closing on its own. Reading the page is mostly solved. Trusting what comes off it is where the work is now.
Let us handle that part. We can't wait to see what you build on top!


