Lovelaice.June 2026
— Lovelaice case study

52% to 88%
accuracy.
Three weeks.

We came into a live AI feature without a quality measurement baseline, alongside a list of improvements that had sat on the roadmap for months. We delivered measurable quality uplift of 36% vs their current baseline, on the model they were already running.

Vertical
Industrial SaaS Series A
Engagement
Eval setup & optimization
Their time investment
~120 minutes total
— Who they are

A live AI feature,
without a measurement baseline.

A Series A B2B software company serving industrial customers. Their AI feature ingests source documents from their customers and converts them into structured JSON that end users execute inside the UI.

Live in production. Inputs span multiple languages and carry dense, domain-specific vocabulary alongside complex visual elements.

Where the gaps were

The team could see that the output had room to improve. End customers were manually editing the AI-extracted data inside the UI before passing it downstream. Improvements were on the roadmap for the quarter, but no clear action plan. Without a baseline, each prompt change was hard to evaluate.

If you're running an AI feature in production, some of these may sound familiar:

  • No measured baseline yet. No single accuracy number — only subjective quality assessment.
  • Quality discussed as one broad concept rather than a set of named, fixable failure modes.
  • The default lever is trying a newer model, even when the gains are somewhere else.
  • The rarer input formats fail in ways no one catches until a customer reports it.
— What it cost them
Customers were cleaning up almost every document by hand.
Each manual edit was a small tax on every document processed. Without a measurement in place, the team had no way to tell which prompt changes actually reduced that tax, and which ones quietly made it worse.
— The engagement

What Lovelaice did.

Three weeks, ~120 minutes of their team's time, seven steps:

  1. 1.Ingested 18 real, very complex (10–100+ pages) documents spanning their full format mix.
  2. 2.Ran a 30-minute working session with their CTO and PM to name what “good” actually looked like and embed their expertise.
  3. 3.Turned that into a 20-point automatic eval rubric. Each point captured one precise, named failure mode. The team could finally answer questions like “how well does the AI feature manage required sections?” with a number.
  4. 4.Ran a baseline on their production model and prompt against the new rubric. Result: 52% accuracy. The first concrete number for the status quo.
  5. 5.8 iterations across 7 prompt versions, each targeting named failure patterns.
  6. 6.Benchmarked 18 models across 4 providers (OpenAI, Azure, Anthropic, Gemini) on the same 18-document dataset. The most expensive model in the field cost ~200× more per document than the cheapest. Several of the cheapest models landed within 2pp of the winner on accuracy.
  7. 7.60-minute review of the full report with their team.
— The result

Before and after.

+36pp
Accuracy lift. 52% to 88%, on the same model family.
5
Leading models within 2% of the winner. Pick by cost or latency.
~20
20-point automatic eval rubric, running on every future change.

← All case studiesLovelaice · 2026