Lovelaice.June 2026
— Lovelaice case study

18% to 93%
accuracy.
25 experiments.

A sustainability-tech team had a live AI screening agent that all five top-tier model providers failed equally. We helped them stop chasing the model and start measuring what “good” looked like. They shipped a production-ready agent on GPT-5 Mini at 93% accuracy.

Vertical
Sustainability / ESG Early-stage
Engagement
Eval setup & optimization
Their time investment
~120 minutes total
— Who they are

A live AI screening agent,
stuck below 20% accuracy.

A sustainability-tech company building an automated business-screening tool. Their AI agent reads public information about a business and decides whether it meets a defined set of sustainability criteria. The end customer is a financial institution or a procurement team that needs to filter thousands of businesses against ESG rules.

Live in production. Inputs span multi-source unstructured web content per business, with a structured rubric the agent has to apply to deliver a yes/no decision the customer can defend.

Where the gaps were

The team knew accuracy was the problem. They had tried five different model providers — Anthropic, OpenAI, Gemini, DeepSeek, Perplexity. None of them moved the needle. The most expensive run cost $74 for 226 calls and scored 0% reliability. The cheapest cost less than a cent and scored the same. They were on iteration 24 of model selection with no measured progress.

If you have shipped an AI feature that “kind of works,” some of this will sound familiar:

  • No measured baseline. An accuracy number was missing, and “let's try a newer model” was the default lever.
  • Quality treated as one fuzzy concept rather than a set of named, gradable criteria.
  • Cross-provider iteration without a control. Each new model was a hope rather than a comparison.
  • Source documents and evaluation criteria were drifting apart, so even a perfect model would have hit the same ceiling.
— The real cost
Every “wrong” answer hit the end customer’s decision pipeline.
An ESG screening tool that is wrong 80% of the time is not a tool the financial customer can defend. The product was live, but the team could not honestly recommend it at scale.
— The engagement

What Lovelaice did.

Three weeks, ~120 minutes of their team's time, seven steps:

  1. 1.Ran the baseline experiment on their production setup. Result: 18% accuracy across all five model providers. The first measured number anyone on the team could point to.
  2. 2.Held a 30-minute working session with their product and domain leads to name what “correct” actually meant on a per-criterion basis. Found two sources of ambiguity in the ESG criteria themselves that no model could have resolved.
  3. 3.Rebuilt the evaluation rubric against the cleaned-up criteria. Re-labelled a focused test set so two domain experts agreed on every row before any model was retested.
  4. 4.Ran a targeted prompt iteration loop. The hypothesis was no longer “find a better model” — it was “find the prompt structure that lets the strongest model actually do the task.”
  5. 5.Benchmarked the new prompt across the same five providers, then narrowed to three. Cost per call dropped from $0.12 on the best earlier variant to a fraction of a cent on the eventual winner.
  6. 6.Found the production winner: GPT-5 Mini at 93% accuracy. Not GPT-5. Not Claude Sonnet. The cheaper “one grade down” model — once the task was framed correctly.
  7. 7.60-minute review with their team. Production deployment plan in hand at the end of the meeting.
— The result

Before and after.

+75pp
Accuracy lift. 18% to 93%, on a cheaper model than they had been testing.
5 → 3
Provider field narrowed to the three worth a serious test. Two ruled out by data, not gut.
~100×
Cost-per-call gap from worst to winning model. The winner was on the cheap end.

← All case studiesLovelaice · 2026