— Who they are
A golden set, reviewed once.
Then checked by hand.
A healthtech company with a consumer-facing AI feature. Each user fills in a detailed health profile, and the feature returns a personalised set of recommendations drawn from the company's own catalogue. Live, in front of real people making real health decisions.
Where the gaps were
They weren't flying blind. They had a golden dataset — real profiles paired with the recommendations an expert had validated. But that review happened once, externally, and the feedback arrived unstructured. It never became something the team could run against, again and again. Day to day, quality was spot-checked one recommendation at a time, by eye.
It's a familiar shape. Some of these may resonate:
- A golden set reviewed once and never run systematically — only spot-checked, case by case, by a person.
- Expert feedback delivered unstructured, so it never compounded from one iteration to the next.
- No measure of reliability. Run-to-run consistency on the same profile was never measured. Nobody knew whether the feature held up on a second try.
- Quality treated as one broad concept, not named failure modes — and model choice made on cost and latency, never benchmarked on their own data.
— The real cost
One-by-one review can’t see a pattern.
A reviewer checking outputs one at a time sees a series of plausible answers. They have no way to see what's being systematically dropped, or skewed, across all of them. The errors that matter most are the ones that only exist in aggregate.
— The engagement
We turned a one-off review
into a system.
The golden set they already trusted, run at scale — with every signal compounding in one place. Seven steps:
- 1.Took their validated golden set and made it the foundation — real profiles, each with the expert's expected output, now something they could test against on demand.
- 2.Ran it at scale: the same profiles through multiple models, repeatedly, every run scored automatically against the expected output. Not a spot-check — the whole set, over and over.
- 3.Compounded the signals. Expert judgement, automatic metric checks, and user feedback captured together and carried forward, so each iteration built on the last instead of starting from scratch.
- 4.Turned “quality” into a set of named, automatic checks — schema validity, recommendation count, dosage formatting, expert-aligned selection, catalogue compliance, a safety gate, and output integrity.
- 5.Ran error analysis across every run and every failure — looking at all the misses together, not one at a time.
- 6.Surfaced a systematic bias the team wasn't looking for, and named the distinct failure categories behind it.
- 7.Iterated fast against those named failures — and, because the loop was now systematic, found the model-and-prompt combination that hit their target at the best cost and latency.
— The reveal
What one-by-one checking can't see.
50%
Expert-aligned, at scale. Across every run, the feature matched the expert about half the time. Invisible to spot-checks.
91%
Broad quality score that masked the gap. Healthy on the surface, half-right underneath.
6
Named failure modes, each with a check behind it. Quality stops being one vague concept.
| What we measured | What we found |
|---|
| Broad quality score, incumbent setup | 91% · looked healthy |
| Expert-aligned selection, at scale | 50% · the gap one-by-one checks miss |
| Reliability across repeated runs | Unmeasured — scored on every run |
| Named, measurable failure modes | 0 — 6 categories, each with a check behind it |
| Models benchmarked on their data | 1 — multiple, on the same golden set |
| The golden set itself | Reviewed once, by hand — run at scale, on demand |
— The aggregate was lying
A healthy headline number hid a coin-flip on the only thing that matters.
Surface checks were near-perfect, and they made the feature look done. Expert-alignment is the metric that says whether a recommendation is actually right for the person — and at scale it sat at 50%. You only see that if you run the golden set on purpose, over and over.
— The unlock
A pattern no one was looking for.
Running the golden set at scale let us do what a human reviewer can't: look at every failure together. The selection errors weren't random. They were systematic — the model consistently favoured some items and suppressed others, independent of the user in front of it.
| Behaviour across mismatched cases | Frequency |
|---|
| One item — systematically omitted, regardless of profile | missing in 14 of 16 |
| A second item — inserted regardless of profile | added in 10 of 16 |
| A third item — inserted regardless of profile | added in 9 of 16 |
A reviewer spot-checking a handful of outputs sees a handful of plausible answers. They have no way to see that one item is dropped on nearly every one, and others pushed in where they don't belong. The pattern only exists across runs — which means manual review, however expert, is structurally blind to it. The team wasn't aware of this bias. They weren't looking for it. It surfaced only because the experimentation was systematic.
— Why it matters
A systematic bias is the same miss, every time. That makes it actionable.
A systematic skew points at something concrete — the retrieval step, the catalogue, or the prompt itself — instead of leaving the team with “the model is a bit off.” In a regulated space, a named failure is the kind your team is accountable for fixing, and the kind the auditors expect you to have traced.
— The payoff
When iteration is systematic,
the gains compound.
With a systematic loop in place, improvement stopped being guesswork. The same infrastructure that surfaced the bias also showed that the model they'd chosen — picked on cost and latency, never tested on quality — wasn't the best option on their own data.
| Before | After |
|---|
| Model chosen on | Cost & latency, untested on quality | Measured quality on their own data |
| The chosen model | Incumbent frontier model | A smaller, lower-cost, faster model |
| Iteration loop | One-off review, unstructured feedback | Systematic, compounding across runs |
| Expert-aligned selection | 50% baseline | Improved against target with a new prompt |
— Better, cheaper, faster — provably
The right model wasn’t the one with the reputation. It was the one that measured best on their data.
Lower cost and latency than the incumbent, and higher on the metric that matters once the prompt was fixed — a decision they can defend with evidence, and re-check automatically on every future change.
— Deliverables
What they walked away with.
- Their golden set turned into a living test — run at scale, on demand, instead of reviewed once by hand.
- A system where expert judgement, automatic checks, and user feedback compound from one iteration to the next.
- A set of named, automatic failure modes — including the systematic bias one-by-one review can't surface.
- Reliability they can measure: the same set, across models, run after run.
- A model benchmark on their own data, and an improved prompt that moved the metric that matters toward target.
- The ability to catch the next unknown failure before a user does — not after.
— Why this matters in health
Bad output here doesn’t look like an error. It looks like a clean, plausible plan a real person acts on.
The risk in consumer health AI isn't a crash on screen — it's a confident, well-formatted recommendation a user follows because nothing told them not to. Health-adjacent AI also sits in the part of the EU AI Act that expects demonstrable, ongoing quality controls. A system that compounds evidence across every run is exactly the kind of audit trail those controls require.
— The math
What this would have costed in-house.
Turning a one-off review into a systematic loop, running the golden set at scale across models, and surfacing an aggregate bias no reviewer could see one at a time. Here is the time that would have taken in-house, against the few weeks and small slice of their team's time we delivered it in.
— Important context
In practice, this path wasn't available to them. The systematic aggregation across runs, the metric decomposition, and the multi-model benchmark all require infrastructure the team didn't have. The numbers below assume they could hire or borrow that capacity internally.
| Scenario | Engineering time | PM / Expert time | Calendar time | vs. Lovelaice |
|---|
Best case focused engineer, no interrupts | ~12 days infra + iteration | minimal | ~4–6 weeks focused sprint | ~2× longer, unbounded team time |
Realistic stop-start, no eval infrastructure | ~25 days infra + iteration + benchmark | ~12 days expert review cycles | ~3–4 months stop-start work | ~6× longer, ~180× more team hours |
Lovelaice eval infrastructure, on their behalf | handled | ~120 minutes total 1 kickoff + 1 review | a few weeks end-to-end | baseline |
Best case. A focused AI engineer with eval methodology in hand could build the systematic loop, run the golden set at scale, and land a new model-and-prompt combination in roughly 12 days of dedicated work, which translates into a 4–6 week sprint once you factor in reviews and calendar friction. And at the end of that sprint the aggregate bias that surfaced here may still be invisible — that pattern only shows up when you look across every failure together.
Realistic. Teams without eval infrastructure fall back on the loop they already have: expert reviews of a handful of outputs, followed by a model swap and a re-check by eye. In our experience each cycle absorbs ~3 engineering days plus ~1–2 expert-review days, and stretches across ~2 calendar weeks. Reaching this team's outcome that way takes ~37 person-days of team effort, spread over ~3–4 months of stop-start work. And a systematic bias may still not surface, because manual review is structurally blind to it.
Lovelaice. A few weeks end-to-end. ~120 minutes of the team's time. The golden-set-at-scale infrastructure, the metric decomposition, the multi-model benchmark, and the aggregate-error analysis — all happen outside their calendar.
Neither in-house figure counts the risk of a confident-looking recommendation reaching a real user, or the audit-trail work required by health-adjacent AI regulation that a compounding system produces as a by-product.
— Where this fits
If any of these feel familiar.
- You have a golden set or expected outputs, but you only ever check them one by one
- Your expert feedback arrives unstructured and doesn't compound from one iteration to the next
- You can't say how reliable your feature is — whether the same input holds up across runs
- You suspect there's a pattern or a bias in the failures, but no way to see it in aggregate
- You picked your model on cost and latency and never benchmarked quality on your own data
- You're in a regulated or health-adjacent space where “it looked fine” isn't a defensible standard
Then the same engagement applies. We turn your golden set into a system — run at scale across models, with expert insight, automatic checks, and feedback compounding in one place — and surface the patterns one-by-one review can't. You keep all of it.
— Next step
Book a demo, and we will scope your engagement against the same playbook.
Anonymised reference available on request.
— Related case studies
More results from real projects.