Reading messy invoices: why extraction accuracy is a distribution, not a number
Every document-AI vendor quotes an extraction accuracy figure, usually somewhere in the high nineties. The number is comforting and nearly meaningless on its own, because accuracy is not one value — it is a distribution across invoice types, and the average tells you nothing about the tail.
Where the average comes from
Clean, digital invoices from a familiar vendor extract near-perfectly. Average those with a few hard cases and you get a flattering headline. But your team does not spend its time on the easy invoices; it spends its time on the scanned, the handwritten, the multi-page, the foreign-format, and the just-plain-wrong. The tail is the work.
The questions that matter
- What is the accuracy on scanned and low-quality images, not just digital PDFs?
- How does the system behave on invoice formats it has never seen?
- When it is unsure, does it flag the field for review or silently guess?
Why "knows what it does not know" beats raw accuracy
A system that is 98% accurate and confidently wrong on the other 2% is dangerous, because the errors flow straight to the ledger. A system that is 96% accurate but flags its uncertain extractions for a human is safer and, in practice, more useful. Calibrated confidence — knowing when to ask — matters more than a slightly higher headline number.
When you evaluate extraction, bring your ugliest invoices, not your cleanest. The cleanest ones every vendor handles. The ugly ones are where the real product lives.