Advanced AI evaluation
For teams past the basics. Deterministic metrics, error analysis, agent harness evaluation and the failure modes that only show up when you're serving real traffic. Assumes you already have an evaluation loop and want to make it hold up under scrutiny.

Article
If the harness is the moat, evaluation is how you defend it

Article
Temperature, max tokens, and streaming: the LLM settings that quietly change your output

Article
Error analysis for AI: turning messy review notes into a fix list

Article