Labelled data you can defend in a review.
Human annotation for systems that have to be right. Domain specialists work to a written spec, disputed items are adjudicated rather than averaged, and every batch ships with its agreement rate and its edge cases named.
Unlabelled
A batch arrives as items with no structure: transactions, calls, notes, photographs. Nothing here is wrong yet, and nothing here can train anything.
Three passes
Every item is labelled three times, by specialists who cannot see each other’s work. A coarse shape appears. Most of the batch settles immediately.
Disagreement
The items that split the passes are not noise. They sit on the boundary between classes, which is precisely where a model in production gets things wrong.
Adjudicated
A named specialist rules on each contested item and the spec is amended underneath it. The boundary stops moving, and it stays where it is for the rest of the batch.
Specialists, screened on your task
Annotators qualify on your data with your guidelines, and we publish who labelled what.
Adjudicated, not averaged
Multi-pass review with a named adjudicator on disagreements, so the label has a reason behind it.
Reported, batch by batch
Agreement rate, throughput and the edge cases that forced a spec change, in writing.
What comes with it
Three passes, one ruling, and the reason in writing.
Where annotators disagree, most vendors take the majority and move on. We open the item, rule on it, and record why. Every ruling is attributable, and the ones that expose a gap amend the spec for the rest of the batch.
A reversal flag records an instruction, not an outcome. Until the ledger settles at value date the customer is still out of funds, and a model trained on “reversed” will close the ticket early.
One agreement number hides the class that will break you.
We report agreement per class, and the pairs that annotators confuse most. Those pairs are where a model fails in production, and they are the pairs your reviewers should read first.
What ships with every batch
Labels are the smallest part of the delivery. The rest is the evidence that lets your model risk team sign them off.
Cohen’s κ per class and per annotator pair, the confusion pairs ranked by volume, and drift against the previous batch. The weakest class is named rather than averaged away.
Every overturned label with the original three passes kept intact, the named adjudicator, and the reasoning. Re-labelled items stay flagged in place.
The items that forced a spec amendment, dated, with worked examples. Held back from training and handed over as an evaluation set.
Who labelled what and when, the access granted to each annotator, and the environment the data never left. Signed off before the batch closes.
The rest of the system
Own your own AI future.
Label the data your model will be judged on, with the evidence attached.