Lovelaice.June 2026
— Lovelaice case study

The failure
no one was
looking for.

They did everything right. A validated golden dataset, reviewed by an expert, checked case by case. What they couldn't do by hand was run it systematically — the same profiles, across models, again and again — and look at every failure together. When we did, a consistent bias surfaced: one the team wasn't aware of, wasn't looking for, and would never have found one recommendation at a time.

Vertical
Healthtech Consumer health
Engagement
Eval audit & optimization
Approach
Systematic, multi-run pattern analysis
— Who they are

A golden set, reviewed once.
Then checked by hand.

A healthtech company with a consumer-facing AI feature. Each user fills in a detailed health profile, and the feature returns a personalised set of recommendations drawn from the company's own catalogue. Live, in front of real people making real health decisions.

Where the gaps were

They weren't flying blind. They had a golden dataset — real profiles paired with the recommendations an expert had validated. But that review happened once, externally, and the feedback arrived unstructured. It never became something the team could run against, again and again. Day to day, quality was spot-checked one recommendation at a time, by eye.

It's a familiar shape. Some of these may resonate:

  • A golden set reviewed once and never run systematically — only spot-checked, case by case, by a person.
  • Expert feedback delivered unstructured, so it never compounded from one iteration to the next.
  • No measure of reliability. Run-to-run consistency on the same profile was never measured. Nobody knew whether the feature held up on a second try.
  • Quality treated as one broad concept, not named failure modes — and model choice made on cost and latency, never benchmarked on their own data.
— The real cost
One-by-one review can’t see a pattern.
A reviewer checking outputs one at a time sees a series of plausible answers. They have no way to see what's being systematically dropped, or skewed, across all of them. The errors that matter most are the ones that only exist in aggregate.
— The engagement

We turned a one-off review
into a system.

The golden set they already trusted, run at scale — with every signal compounding in one place. Seven steps:

  1. 1.Took their validated golden set and made it the foundation — real profiles, each with the expert's expected output, now something they could test against on demand.
  2. 2.Ran it at scale: the same profiles through multiple models, repeatedly, every run scored automatically against the expected output. Not a spot-check — the whole set, over and over.
  3. 3.Compounded the signals. Expert judgement, automatic metric checks, and user feedback captured together and carried forward, so each iteration built on the last instead of starting from scratch.
  4. 4.Turned “quality” into a set of named, automatic checks — schema validity, recommendation count, dosage formatting, expert-aligned selection, catalogue compliance, a safety gate, and output integrity.
  5. 5.Ran error analysis across every run and every failure — looking at all the misses together, not one at a time.
  6. 6.Surfaced a systematic bias the team wasn't looking for, and named the distinct failure categories behind it.
  7. 7.Iterated fast against those named failures — and, because the loop was now systematic, found the model-and-prompt combination that hit their target at the best cost and latency.
— The reveal

What one-by-one checking can't see.

50%
Expert-aligned, at scale. Across every run, the feature matched the expert about half the time. Invisible to spot-checks.
91%
Broad quality score that masked the gap. Healthy on the surface, half-right underneath.
6
Named failure modes, each with a check behind it. Quality stops being one vague concept.

← All case studiesLovelaice · 2026