The 99% Lie: Why the Accuracy Number Every Vendor Quotes Isn't the One That Costs You Money
How accurate is document AI, really? 99% per field on a forty-field invoice can mean a third of your invoices still need a human. The number that matters is the one almost nobody quotes.
A vendor tells you their extraction is 99% accurate. Your invoice has forty fields — supplier, number, date, a dozen line items with quantities and prices, three tax lines, a subtotal, a total, payment terms. Do the sum the vendor didn't.
If every field has a one-in-a-hundred chance of being wrong, the chance that all forty are right is 0.99 multiplied by itself forty times: 0.669. Two invoices in three come through clean. One in three arrives with at least one error on it.At 99.5% per field it's one in five. At 99.9% — a number nobody honestly offers on real documents — it's still one in twenty-five.
Now do the sum that actually matters. Somebody has to find that third. If the system can't say which invoices are the flawed ones, a person opens all of them, or none. If it can, a person opens only those, finds the wrong field, fixes it and moves on. Either way the cost is a person's time per document. The 99% was never about them.
The vendor's number is true. The lie is in the unit. Accuracy is quoted per field. You pay per document. That gap is the whole of this piece.
One caveat: that sum treats every field as its own coin toss, and real errors bunch up in the difficult documents. The first objection below does the corrected sum. It makes the 99% look worse, not better.
Field accuracy vs document accuracy vs straight-through rate(what "accuracy" is actually counting)
There are four numbers hiding behind the word, and they describe different things.
Field accuracy — of every field the system read, what share matched the truth. One number per box on the form, averaged across thousands of boxes. This is the 99%. It describes the model.
Document accuracy — of every document, what share had every field right. This is the compounding sum from the top of the piece. It is always lower than field accuracy, and the gap widens with every field on the form.
Straight-through rate— of every document, what share went through with no human touching it. Finance and operations teams know it as straight-through processing, or STP. This is the one that maps to hours and money. A pipeline that stops when it isn't sure has a lower straight-through rate than one that guesses — and that's the point.
False-pass rate— of the documents that went straight through, what share were wrong. This is the number that keeps straight-through rate honest. A vendor can reach 100% straight-through by passing everything; the false-pass rate is what they'd have to confess.
The first two describe the model. The last two describe your Tuesday.
The public example: Amazon counted reviews per thousand, not accuracy
Does anyone actually run an automation on these numbers? The clearest public example isn't a document company at all. It's Amazon — a different task, and a harder one: following shoppers round a shop is a tougher vision problem than reading an invoice. The unit is the same.
Just Walk Out opened to the public in 2018: cameras and sensors track what shoppers pick up, and the receipt follows them out of the door. The Information first reported in 2023 — and repeated when the Fresh rollback was announced in April 2024 — that as of 2022 about 700 of every 1,000 transactions had needed human review — against an internal target of fewer than 50 — with more than a thousand people in India reviewing and labelling footage.
Amazon disputed the report. A spokesperson called the idea of reviewers watching shoppers live "misleading and inaccurate." Jon Jenkins, the vice-president then running the product, told Axios that the India team was "way less than 1,000" people, that they reviewed some video after the fact to train the models, and that they helped verify receipts in "a small percentage of cases."
Here is what isn't disputed. The same month, Amazon began removing Just Walk Out from its US Fresh grocery stores in favour of Dash Carts — smart trolleys that scan as you shop — while keeping it at airports, stadiums and third-party venues.
We don't know Amazon's real rate, and neither does anyone outside Amazon. What's on the record is that the company hada target — fewer than five transactions in a hundred reviewed — and pulled the product from the stores where it mattered most. The point isn't that the vision system was bad. It's that Amazon budgeted in the right unit, reviews per thousand transactions, while the market only ever heard "just walk out." The headline and the operating metric were different numbers. They always are.
So let's name it: the Field Fallacy
The Field Fallacy is quoting accuracy in the unit that flatters the model and budgeting in the unit that pays the people. Almost every vendor page in this category does it — ours has too — and almost every buyer accepts it, because 99% sounds like a finished conversation.
Two vendors, both "99% accurate"
Same 1,000 invoices a day. Same forty fields. Same quoted number. The figures below are a worked example, not anyone's real results — but the arithmetic in them is exact.
Read across the first two rows: identical. Read down: the only rows that predict your cost are the two nobody quoted. And Vendor B — the one with the better straight-through rate, the one whose demo looks faster — is the one that costs more, because the 230 invoices it got wrong didn't disappear. They went into the ledger.
The 99% told you nothing. The pair of numbers underneath told you everything.
Do the sum on your own documents
Here is the one line of arithmetic. Documents a day × (1 − field accuracyfields) × minutes to fix one ÷ 60 = hours a day. On the example above, at four minutes a document: 1,000 × 0.33 × 4 ÷ 60 is about 22 hours a day — nearly three people — from a system quoted at 99%.
What a pilot report should contain(and how AutomatR produces it)
A pilot that reports field accuracy has told you how good the model is on your documents. Fine. It hasn't told you what the process will cost to run. The report worth signing off on has four parts, all about the same set of documents. This is the report we write.
What the documents were, how many, from how many sources, over what period — and what share were the ugly kind. Your set, not a sample set of the vendor's. A pilot on clean samples is a demo.
The first is the number you were quoted. The second is what it comes to on your forms, with your field counts — measured directly, by counting the documents in the set that came through with every field right, not inferred from the first. Seeing them together is usually the moment the room goes quiet.
How many documents no one touched, and how many of those were wrong. Ask where the false-pass figure came from. Ours comes from two places: a random sample of passed documents checked by hand, and reconciliation flags downstream. Every document that stopped says why — a field below threshold, a total that doesn't reconcile, a supplier that matched two records — so the queue is a list of findings, not failures. And the thresholds that decide what stops are set with you, per field and per document type: the trade between straight-through and false-pass is your business decision, and the report shows where the dial is set and what moving it does to both numbers.
Exceptions per hundred documents, minutes per exception. This is the line the CFO reads, and it's derived from the two rates above, not from the 99%.
A document reaches straight-through only by clearing three gates:
- Confidence. Each extracted field carries a confidence score from the extraction model; the document passes only when every required field clears its threshold.
- Cross-field reconciliation. Independently of confidence, checks run on the document as a whole — line items to the subtotal, tax to the rate, subtotal plus tax to the total. Any inconsistency stops the document, however confident the individual fields were.
- Master-data match.Extracted identifiers are matched against the customer's master data and must return exactly one record.
False-pass is estimated two ways: a random sample of passed documents is checked by hand during the pilot, and downstream reconciliation — payment mismatches, supplier queries — is fed back as a signal. The AI designs the workflow; a deterministic engine runs it, which is why every one of those decisions is logged with its reason.
You'll have noticed there is no AutomatR figure in this post. A rate is only worth quoting with its document set beside it, and the set that matters is yours. Send us five of your ugliest — the rotated scan, the three layouts from one supplier, the handwritten field, the forty-page pack, the spreadsheet nobody trusts — through automatr.tech/contact and the report will have both rates, with the set they came from.
Five objections
"Field errors aren't independent. Isn't your 67% wrong?"
It's the pessimistic end, yes. Errors cluster. A skewed scan or an unfamiliar layout gets six fields wrong at once; a clean PDF from a regular supplier gets none. So on real documents more than 67% come through clean, and the table's 330 moves with it.
Now look at what that does to the 99%. A thousand invoices is 40,000 fields; at 99%, that's 400 errors. Say they fall on 150 documents: 85% of your invoices are clean — but those 150 are being read at about 93% per field, not 99. The average was hiding two populations: documents the system reads almost perfectly, and the ugly ones from The Data-Readiness Myth, read far worse than the brochure says. Clustering changes how many documents stop. It doesn't change the question: how many stopped, and how many of the ones that didn't were wrong?
"What if our vendor's 99% is document-level?"
Then you have the rarer, better number, and it deserves two follow-ups. First: what counted as a document, a one-page invoice or a forty-page pack? The unit still matters. Second: what's the false-pass rate that goes with it? Document-level accuracy tells you how often the model was right. It still doesn't tell you how often it was right and nobody checked.
"Humans aren't 99% either, are they?"
True, and it cuts the other way. A person who is unsure stops and asks. The question for any system isn't whether it beats a human on a good day; it's whether it stops when it should. That is exactly what the false-pass rate measures, and it's why a lower straight-through rate can be the better product.
"Can't straight-through rate be gamed?"
It can — by passing everything. Which is why it never travels alone. Straight-through with false-pass beside it can't be gamed: lower the bar to raise the first, and the second rises with it. A vendor who can answer the first and not the second hasn't really measured the first.
"Wouldn't a vendor say this?"
Yes. So take the Two-Number Card below to our demo before anyone else's. Bring those five ugly documents and ask for the two numbers on those.
The Two-Number Card
Take these to any vendor demo, including ours. Start by settling the unit — per field, or per document? If per field: how many fields on my document, and what does that compound to? Then ask for the two numbers.
- 01Of the last hundred documents, how many did no human touch?That’s the straight-through rate. On which documents?
- 02Of those, how many were wrong?That’s the false-pass rate. How was it measured?
If a vendor can quote you an accuracy number faster than they can answer the second question, you've learned which one they measure.
Of the last hundred documents your automation processed, how many did no human touch — and does anyone in the building know?
Further reading: going deeper on accuracy
- →The Data-Readiness Myth — where the ugly third comes from, and the five documents to bringautomatr.tech/blog/data-readiness-myth
- →The Four Surfaces — the same two rates, plus the third question that goes with them, in contextautomatr.tech/blog/four-surfaces
- →The Four Exits — what the engine is allowed to do with what it readsautomatr.tech/blog/four-exits-agent-data-leaks
- →Axios, “Amazon pushes back on perception of Just Walk Out” (17 April 2024) — Amazon's account, in its own wordsaxios.com
- →Fortune, “Amazon’s co-inventor of Just Walk Out sets the record straight” (17 April 2024)fortune.com