We ran the generic LLM judge templates from all three tools against 34 human-labeled AI responses, across 10 models, twice. Same dataset, same protocol, three vendors. This is what we found.

Every AI observability vendor now ships a set of "quality judge" templates. Helpfulness. Relevance. Correctness. Hallucination. Drop them into your pipeline and get scores back. Product teams adopt them because they arrive already written, they look like measurement, and nobody wants to be the person who couldn't produce an accuracy number for the AI feature.

We wanted to know which vendor's templates were the strongest. So we put them side by side.

Over 5,000 verdicts later, the answer is that the vendor barely matters. All three vendors' generic templates cluster inside a 6-point band, well within run-to-run noise. And every one of them fails the same way, in one of two directions, determined by what the template asks, not by which vendor wrote it.

Thesis

The vendor is not the variable. The rubric is.

The full experiment (8 judges, 10 models, 2 runs, 34 real agent responses) is written up in the battle of the LLM judges. This piece pulls the vendor-vs-vendor slice out of it, because the comparison keeps coming up in evaluation-tool decisions and we've never seen anyone publish the numbers.

Enjoyed reading? Add us as a preferred source in Google to support us.

Add Lovelaice as a preferred source