Skip to main content
All articles

Data Extraction Tips

Confidence scores only matter if someone acts on them

A score attached to a field is not accuracy. It is a routing instruction, and most extraction projects never wire up the route.

Written by
Marcus Feld
Extraction Lead, Datacoll8
Published
Reading time
8 minutes

Every serious extraction system attaches a confidence score to each field it reads. Very few organisations do anything with it. The score is displayed in a dashboard, admired during procurement, and then every extracted value is written into the record regardless of what the number said. At that point the score is decoration.

A confidence score is not a claim about accuracy. It is the model saying how much of its own output it would defend. Treated properly, it is a routing instruction: this value goes straight through, that value needs a human.

Set the threshold per field, not per document

A single global threshold forces the same tolerance onto values with wildly different consequences. Getting a counterparty address slightly wrong is an inconvenience. Getting a renewal date wrong can lose a client a right they did not know they were about to give up. Those two fields should not share a threshold.

Group your fields by what a mistake costs. Fields where an error is caught immediately by someone reading the document can sit at a relatively permissive threshold. Fields that drive a diary entry, a payment or a deadline should be set tight enough that the queue is used, and in some cases should be routed to review unconditionally.

Show the reviewer the page, not just the value

The most common design mistake in verification is presenting a list of low-confidence values as a spreadsheet. The reviewer then has to open the source document, find the clause, and rebuild the context the model already had. That is slower than reading the document from scratch, and reviewers respond to it exactly as you would expect: by approving in bulk.

Show the extracted value beside the region of the page it was read from. The decision becomes a glance instead of an investigation, review time per field falls to a few seconds, and, importantly, the approvals you get back mean something.

A review queue that is tedious will be cleared, not read. Design for the glance and you get genuine verification.
Marcus Feld, Extraction Lead

Log corrections as training signal, not as errors

When a reviewer changes a value, that is the most valuable data your pipeline will produce all week. It tells you exactly which document type, which clause and which layout the model is struggling with. Captured properly, a month of corrections tells you where to retune, which templates to split and which counterparty paper needs its own handling.

Firms that treat corrections purely as a quality complaint learn nothing from them. Firms that review correction patterns on a cycle watch their queue volume fall, because they are fixing causes rather than instances.

Watch the ratio, not the headline number

  • Auto-accept rate: the share of fields that clear the threshold without review
  • Correction rate within the queue: how often a reviewer actually changes a flagged value
  • Escaped error rate: errors found downstream that never entered the queue
  • Time per reviewed field: the honest measure of whether your interface works

A high auto-accept rate with a low escaped error rate is the goal. A high auto-accept rate with a rising escaped error rate means your thresholds are too loose and you are shipping mistakes. A low correction rate inside the queue means the opposite: you are sending humans work they did not need to see, and you can tighten up.

Those four numbers are worth more than any accuracy percentage on a datasheet, because they describe your documents, your thresholds and your reviewers rather than a vendor benchmark set. Instrument them first, and the tuning work tells you where to go.

Written byMarcus FeldExtraction Lead, Datacoll8

Book a demo

See it run against your own documents

Bring three or four representative agreements to the demo. We will run them through classification, extraction and verification live, and tell you plainly where the pipeline would need tuning for your paper.

Demos run Monday to Friday, 08:30 to 18:00 GMT