Reading messy invoices: why extraction accuracy is a distribution, not a number

Every document-AI vendor quotes an extraction accuracy figure, usually somewhere in the high nineties. The number is comforting and nearly meaningless on its own, because accuracy is not one value — it is a distribution across invoice types, and the average tells you nothing about the tail.

Where the average comes from

Clean, digital invoices from a familiar vendor extract near-perfectly. Average those with a few hard cases and you get a flattering headline. But your team does not spend its time on the easy invoices; it spends its time on the scanned, the handwritten, the multi-page, the foreign-format, and the just-plain-wrong. The tail is the work.

The questions that matter

  • What is the accuracy on scanned and low-quality images, not just digital PDFs?
  • How does the system behave on invoice formats it has never seen?
  • When it is unsure, does it flag the field for review or silently guess?

Why "knows what it does not know" beats raw accuracy

A system that is 98% accurate and confidently wrong on the other 2% is dangerous, because the errors flow straight to the ledger. A system that is 96% accurate but flags its uncertain extractions for a human is safer and, in practice, more useful. Calibrated confidence — knowing when to ask — matters more than a slightly higher headline number.

When you evaluate extraction, bring your ugliest invoices, not your cleanest. The cleanest ones every vendor handles. The ugly ones are where the real product lives.

Put it into practice.

See how Astridex automates this on your actual workflows.