Case study
The Edit Rate
A human corrected the model's output every time it got something wrong. Nobody was counting. Turning that correction into a metric is how we found out the model was degrading.
- Edit rate on critical events fell from 60% to 25%
- Previously invisible work measurable down to per-task effort
- A model-degradation signal that moved before anyone complained
- Workflow automations cut processing time 30%
Context
Supio ran a human-in-the-loop annotation operation. Vendor annotators reviewed and corrected AI model output over medical and claims records before anything went out the door. It was a large, continuous, expensive process, and when I arrived not one part of it was measured.
That meant two things were invisible at once. How much human effort the AI was actually costing us, and whether the model was getting better or worse over time. Those sound like different questions. They have the same answer.
How do you know a model is degrading?
This is the question I care most about in AI products, and it is harder in production than it looks in a notebook. Your offline numbers were computed on a test set. Production sends you a different distribution every month, and it doesn't come with labels.
Meanwhile the people best placed to notice were the annotators, who were silently absorbing the problem. When the model got worse, they just edited more. The cost showed up as effort, not as an alert.
The options for a quality signal
Periodic accuracy sampling against a labelled set
PassedRigorous, and it answers a question about the sample rather than about production. It needs labelled truth at volume, which is the exact thing that is expensive here, and it tells you nothing between samples.
Wait for Operations to escalate
PassedThe signal genuinely exists in the room, but it arrives late and as anecdote. By the time a complaint is loud enough to act on, you have paid for months of extra correction and you still can't say when it started.
Instrument the correction itself
ChosenEvery edit an annotator makes is a human judgment that the model was wrong, produced as a by-product of work already being paid for. Edit rate is the inverse of acceptance, it needs no separate labelling exercise, and it moves before anyone thinks to complain.
Why the by-product was the right signal
The ground truth already existed. It was being generated continuously by people whose job was to produce it, and then thrown away because nothing captured it. Measurement problems that look expensive are often like this, where the data is being created and discarded rather than never created at all.
It also meant the metric aligned with the thing the business actually cared about. Edit rate isn't a proxy for cost, it is very close to cost, because every point of it is human minutes. A quality metric that doubles as a cost metric gets looked at by people who would otherwise never open a model dashboard.
What I built
I defined the operational signals worth tracking and partnered with Operations to instrument them, including timestamps inside the annotation interface itself, last-clicked and field-submitted, so we could compute total review time per annotator and per-task effort underneath the quality numbers. Work that had been completely invisible became measurable down to the individual task.
The reporting shipped as a hosted Streamlit suite that teams across several functions self-served for more than six months. I should be straightforward about that piece: Streamlit was the delivery surface and it was built with heavy AI assistance. What I own here is the measurement design and the outcome, not app engineering.
What the signal actually changed
Edit rate on critical events went from 60% to 25%, and it drove three different kinds of decision, which is the part I'd point at rather than the number.
It informed model work. A rising edit rate meant the model was degrading, and that put the question in front of engineering with evidence attached. To be precise about who did what: I surfaced the signal and sat in the retraining decisions, engineering and data science did the retraining and the prompt work.
It drove vendor coaching and staffing. Per-task effort and productivity data let us see which teams were slow where, and size staffing against the workload that was actually coming. I also built fraud detection into the vendor-invoicing reporting, which flagged overstated annotation volumes we were being billed for.
And it pointed the roadmap. Knowing where annotator effort actually sat told us which steps were worth automating rather than guessing. The workflow automations that came out of it cut processing time 30%.
Results
Worth stating the boundary plainly, because it is the first thing a good interviewer will push on. I built and owned the metric, surfaced the degradation, and was in the room for retraining decisions. I did not retrain models or change prompts. This is measurement of a model, not modelling.