Accuracy

The 99% Lie: Why the Accuracy Number Every Vendor Quotes Isn't the One That Costs You Money

How accurate is document AI, really? 99% per field on a forty-field invoice can mean a third of your invoices still need a human. The number that matters is the one almost nobody quotes.

AutomatR TeamOct 5, 202611 min read

A vendor tells you their extraction is 99% accurate. Your invoice has forty fields — supplier, number, date, a dozen line items with quantities and prices, three tax lines, a subtotal, a total, payment terms. Do the sum the vendor didn't.

If every field has a one-in-a-hundred chance of being wrong, the chance that all forty are right is 0.99 multiplied by itself forty times: 0.669. Two invoices in three come through clean. One in three arrives with at least one error on it.At 99.5% per field it's one in five. At 99.9% — a number nobody honestly offers on real documents — it's still one in twenty-five.

Now do the sum that actually matters. Somebody has to find that third. If the system can't say which invoices are the flawed ones, a person opens all of them, or none. If it can, a person opens only those, finds the wrong field, fixes it and moves on. Either way the cost is a person's time per document. The 99% was never about them.

The vendor's number is true. The lie is in the unit. Accuracy is quoted per field. You pay per document. That gap is the whole of this piece.

By the end of this piece you'll be able to turn any accuracy claim into an hours estimate with one line of arithmetic — and know which two numbers to ask for instead.
67%
Documents with zero errors at 99% per field across 40 fields
331
Documents a day, out of 1,000, with at least one wrong field
22 hrs
A day, at four minutes to find and fix each one
SOURCE: (FIELD ACCURACY)^40, COMPUTED · 1,000 × 0.33 × 4 ÷ 60 · ASSUMES ERRORS FALL INDEPENDENTLY — THE FIRST OBJECTION BELOW TAKES THAT ON

One caveat: that sum treats every field as its own coin toss, and real errors bunch up in the difficult documents. The first objection below does the corrected sum. It makes the 99% look worse, not better.

Share of documents with zero errors, by number of fieldsThree curves showing document-level clean rate falling as fields per document rise, at 99%, 99.5% and 99.9% per-field accuracy. At 40 fields and 99% per field, 67% of documents are clean.0%25%50%75%100%020406080100Fields per document40 fields → 67% clean99.9% per field99.5% per field99% per field
DOCUMENT-LEVEL CLEAN RATE = (FIELD ACCURACY)^FIELDS, ASSUMING INDEPENDENT ERRORS. THE VENDOR'S NUMBER IS THE LEGEND; YOURS IS WHEREVER YOUR FORM SITS ON THE X-AXIS.

Field accuracy vs document accuracy vs straight-through rate(what "accuracy" is actually counting)

There are four numbers hiding behind the word, and they describe different things.

Field accuracy — of every field the system read, what share matched the truth. One number per box on the form, averaged across thousands of boxes. This is the 99%. It describes the model.

Document accuracy — of every document, what share had every field right. This is the compounding sum from the top of the piece. It is always lower than field accuracy, and the gap widens with every field on the form.

Straight-through rate— of every document, what share went through with no human touching it. Finance and operations teams know it as straight-through processing, or STP. This is the one that maps to hours and money. A pipeline that stops when it isn't sure has a lower straight-through rate than one that guesses — and that's the point.

False-pass rate— of the documents that went straight through, what share were wrong. This is the number that keeps straight-through rate honest. A vendor can reach 100% straight-through by passing everything; the false-pass rate is what they'd have to confess.

The first two describe the model. The last two describe your Tuesday.

The public example: Amazon counted reviews per thousand, not accuracy

Does anyone actually run an automation on these numbers? The clearest public example isn't a document company at all. It's Amazon — a different task, and a harder one: following shoppers round a shop is a tougher vision problem than reading an invoice. The unit is the same.

Just Walk Out opened to the public in 2018: cameras and sensors track what shoppers pick up, and the receipt follows them out of the door. The Information first reported in 2023 — and repeated when the Fresh rollback was announced in April 2024 — that as of 2022 about 700 of every 1,000 transactions had needed human review — against an internal target of fewer than 50 — with more than a thousand people in India reviewing and labelling footage.

Amazon disputed the report. A spokesperson called the idea of reviewers watching shoppers live "misleading and inaccurate." Jon Jenkins, the vice-president then running the product, told Axios that the India team was "way less than 1,000" people, that they reviewed some video after the fact to train the models, and that they helped verify receipts in "a small percentage of cases."

Here is what isn't disputed. The same month, Amazon began removing Just Walk Out from its US Fresh grocery stores in favour of Dash Carts — smart trolleys that scan as you shop — while keeping it at airports, stadiums and third-party venues.

We don't know Amazon's real rate, and neither does anyone outside Amazon. What's on the record is that the company hada target — fewer than five transactions in a hundred reviewed — and pulled the product from the stores where it mattered most. The point isn't that the vision system was bad. It's that Amazon budgeted in the right unit, reviews per thousand transactions, while the market only ever heard "just walk out." The headline and the operating metric were different numbers. They always are.

So let's name it: the Field Fallacy

The Field Fallacy is quoting accuracy in the unit that flatters the model and budgeting in the unit that pays the people. Almost every vendor page in this category does it — ours has too — and almost every buyer accepts it, because 99% sounds like a finished conversation.

WHAT THE VENDOR MEASURES
Per field, on their sample set, averaged
Honest as far as it goes. Useless for a budget: it says nothing about how many documents a person will open.
WHAT YOU PAY FOR
Per document, on your documents, in human touches
The exception queue, the approvals, the corrections at month-end. Counted in hours, not fields.

Two vendors, both "99% accurate"

Same 1,000 invoices a day. Same forty fields. Same quoted number. The figures below are a worked example, not anyone's real results — but the arithmetic in them is exact.

A worked comparison of two vendors with identical 99% field accuracy, showing how straight-through rate and false-pass rate move in opposite directions and where the cost lands.
Vendor A (stops when unsure)Vendor B (passes anything above a low bar)
Quoted accuracy
99% per field
99% per field
Invoices with at least one wrong field
~330 of 1,000
~330 of 1,000
Straight-through rate60%
600 pass; 400 go to the queue, each with a reason
90%
900 pass; 100 go to the queue
False-pass rate~1.7%
About 10 wrong invoices among the 600 that passed
≥25%
At least 230 of the 330 flawed invoices are among the 900 that passed
Invoices a human touches
400 — in the queue, visible, the same day
100 in the queue — plus ~230 wrong ones nobody flagged
Where the cost lands
Your exception queue
Your ledger, at month-end, or in a supplier's complaint
WORKED EXAMPLE — 1,000 INVOICES, 40 FIELDS, 99% PER FIELD → ~330 CARRY AT LEAST ONE ERROR. A VENDOR THAT PASSES 900 CANNOT HAVE STOPPED MORE THAN 100 OF THEM.

Read across the first two rows: identical. Read down: the only rows that predict your cost are the two nobody quoted. And Vendor B — the one with the better straight-through rate, the one whose demo looks faster — is the one that costs more, because the 230 invoices it got wrong didn't disappear. They went into the ledger.

The 99% told you nothing. The pair of numbers underneath told you everything.

Do the sum on your own documents

Here is the one line of arithmetic. Documents a day × (1 − field accuracyfields) × minutes to fix one ÷ 60 = hours a day. On the example above, at four minutes a document: 1,000 × 0.33 × 4 ÷ 60 is about 22 hours a day — nearly three people — from a system quoted at 99%.

67%documents with zero errors
331documents a day with at least one wrong field
22hours a day spent finding and fixing them
2.8people, at eight hours a day
The hours are a minimum: they assume the system flags exactly the flawed documents and nothing else. The 67% is also a ceiling: any vendor claiming a straight-through rate above it has passed errors. Ask for the false-pass rate before believing any straight-through figure above it. Assumes errors fall independently; see the first objection below.

What a pilot report should contain(and how AutomatR produces it)

A pilot that reports field accuracy has told you how good the model is on your documents. Fine. It hasn't told you what the process will cost to run. The report worth signing off on has four parts, all about the same set of documents. This is the report we write.

01 · The set

What the documents were, how many, from how many sources, over what period — and what share were the ugly kind. Your set, not a sample set of the vendor's. A pilot on clean samples is a demo.

02 · Field accuracy and document accuracy, side by side

The first is the number you were quoted. The second is what it comes to on your forms, with your field counts — measured directly, by counting the documents in the set that came through with every field right, not inferred from the first. Seeing them together is usually the moment the room goes quiet.

03 · Straight-through rate — with the false-pass rate beside it

How many documents no one touched, and how many of those were wrong. Ask where the false-pass figure came from. Ours comes from two places: a random sample of passed documents checked by hand, and reconciliation flags downstream. Every document that stopped says why — a field below threshold, a total that doesn't reconcile, a supplier that matched two records — so the queue is a list of findings, not failures. And the thresholds that decide what stops are set with you, per field and per document type: the trade between straight-through and false-pass is your business decision, and the report shows where the dial is set and what moving it does to both numbers.

04 · The hours

Exceptions per hundred documents, minutes per exception. This is the line the CFO reads, and it's derived from the two rates above, not from the 99%.

For the engineers on the buying committee — everyone else may skip this section

A document reaches straight-through only by clearing three gates:

  1. Confidence. Each extracted field carries a confidence score from the extraction model; the document passes only when every required field clears its threshold.
  2. Cross-field reconciliation. Independently of confidence, checks run on the document as a whole — line items to the subtotal, tax to the rate, subtotal plus tax to the total. Any inconsistency stops the document, however confident the individual fields were.
  3. Master-data match.Extracted identifiers are matched against the customer's master data and must return exactly one record.

False-pass is estimated two ways: a random sample of passed documents is checked by hand during the pilot, and downstream reconciliation — payment mismatches, supplier queries — is fed back as a signal. The AI designs the workflow; a deterministic engine runs it, which is why every one of those decisions is logged with its reason.

You'll have noticed there is no AutomatR figure in this post. A rate is only worth quoting with its document set beside it, and the set that matters is yours. Send us five of your ugliest — the rotated scan, the three layouts from one supplier, the handwritten field, the forty-page pack, the spreadsheet nobody trusts — through automatr.tech/contact and the report will have both rates, with the set they came from.

Five objections

"Field errors aren't independent. Isn't your 67% wrong?"

It's the pessimistic end, yes. Errors cluster. A skewed scan or an unfamiliar layout gets six fields wrong at once; a clean PDF from a regular supplier gets none. So on real documents more than 67% come through clean, and the table's 330 moves with it.

Now look at what that does to the 99%. A thousand invoices is 40,000 fields; at 99%, that's 400 errors. Say they fall on 150 documents: 85% of your invoices are clean — but those 150 are being read at about 93% per field, not 99. The average was hiding two populations: documents the system reads almost perfectly, and the ugly ones from The Data-Readiness Myth, read far worse than the brochure says. Clustering changes how many documents stop. It doesn't change the question: how many stopped, and how many of the ones that didn't were wrong?

"What if our vendor's 99% is document-level?"

Then you have the rarer, better number, and it deserves two follow-ups. First: what counted as a document, a one-page invoice or a forty-page pack? The unit still matters. Second: what's the false-pass rate that goes with it? Document-level accuracy tells you how often the model was right. It still doesn't tell you how often it was right and nobody checked.

"Humans aren't 99% either, are they?"

True, and it cuts the other way. A person who is unsure stops and asks. The question for any system isn't whether it beats a human on a good day; it's whether it stops when it should. That is exactly what the false-pass rate measures, and it's why a lower straight-through rate can be the better product.

"Can't straight-through rate be gamed?"

It can — by passing everything. Which is why it never travels alone. Straight-through with false-pass beside it can't be gamed: lower the bar to raise the first, and the second rises with it. A vendor who can answer the first and not the second hasn't really measured the first.

"Wouldn't a vendor say this?"

Yes. So take the Two-Number Card below to our demo before anyone else's. Bring those five ugly documents and ask for the two numbers on those.

The Two-Number Card

Take these to any vendor demo, including ours. Start by settling the unit — per field, or per document? If per field: how many fields on my document, and what does that compound to? Then ask for the two numbers.

  1. 01Of the last hundred documents, how many did no human touch?That’s the straight-through rate. On which documents?
  2. 02Of those, how many were wrong?That’s the false-pass rate. How was it measured?

If a vendor can quote you an accuracy number faster than they can answer the second question, you've learned which one they measure.

One last question

Of the last hundred documents your automation processed, how many did no human touch — and does anyone in the building know?

Further reading: going deeper on accuracy

SourcesThe compounding arithmetic: document-level clean rate = (field accuracy)^(fields per document), assuming independent field errors; 0.99^40 = 0.669, 0.995^40 = 0.818, 0.999^40 = 0.961. The clustering example: 40,000 fields at 99% is 400 errors; concentrated in 150 documents (6,000 fields) that is a 6.7% field error rate on those documents. The calculator's four-minute default is an assumption for the reader to replace. The Information, May 2023 (first report), restated 2 April 2024 (Tony Hoggett interview and follow-up reporting): more than 1,000 workers in India reviewing and labelling video; about 700 of every 1,000 Just Walk Out transactions requiring human review as of 2022, against an internal target of under 50 per 1,000 — figures Amazon disputes. Amazon's response: spokesperson statement to USA Today calling the live-review framing "misleading and inaccurate"; Jon Jenkins, VP, to Axios (17 April 2024): team "way less than 1,000," reviews video after the fact for training, verifies receipts in "a small percentage of cases." Removal from 28 US Amazon Fresh stores and two Whole Foods in favour of Dash Carts: Bloomberg via Fortune (3 April 2024) and Fortune (17 April 2024); the technology continues at third-party venues and in UK stores. The two-vendor table is a worked example with exact arithmetic, not any company's results. AutomatR's own pilot figures are provided on the customer's documents in the pilot report, not quoted here.
AutomatR
AutomatR TeamAutomatR — Unified Agentic Stack for Enterprises