Skip to content
All case studies

LLM Systems + Evaluation & Cost Optimization

Making Two LLM Systems Measurable, Then Cheaper

Two production LLM systems could report exactly what a run cost and nothing at all about whether the output was any good. I built the missing measurement first, then used it to make the cost reductions safe.

2

production LLM systems moved from cost-only to quality observability

Conformance checks gating every prompt and routing change

Ranking weights replaced with measured values instead of guesses

An unbounded prompt-growth defect found and fixed

The problem

Both systems had precise cost observability and no quality observability whatsoever, which is worse than having neither, because it makes every cost reduction look free. One could tell you the exact cost of a run but not whether the run had produced anything useful, so no prompt or model change could be validated except by a person reading the output and forming an opinion. The other ranked results using weights that had been guessed at and never checked, and its largest available saving would have handed the cheap stage authority over what the expensive stage was even allowed to consider. Nothing in either system could tell you whether shipping that would be free or actively harmful.

What I built

Split evaluation into two kinds of check, because they need different treatment. Conformance checks have unambiguous answers and can block a merge outright. Quality checks are noisy, so they gate on a threshold plus a human signing off, since a hard gate on a noisy metric only teaches people to route around the gate. For the ranking system, built the evaluation set out of labelled outcomes the organization already had and had never used for this purpose, which turned guessed weights into measured ones. Only after that did the cost work start, with every optimization recording before and after in the same harness and stating its quality cost next to its saving.

Technical approach

  • Conformance checks block a merge while quality regressions require a human to sign off, because a noisy metric behind a hard gate trains people to bypass it rather than to fix the regression
  • The evaluation set was built from labelled outcomes that already existed, which is cheaper and more honest than authoring one, provided you exclude the outcomes the old system itself produced
  • That exclusion mattered more than it sounds. Building on output from the system you are replacing quietly invalidates every conclusion downstream, and the invalidation stays invisible for a long time
  • Chose the metric around the failure nobody notices rather than the one that is easy to measure, since a screening system's real catastrophe is the good result that never surfaces at all
  • Cost reductions that route work to a cheaper model were gated on the measured quality of the cheap stage, not on the obvious saving in call volume, because a prefilter that is wrong is invisible in the cost graph
  • Worked out the economics of escalating from cheap to expensive before shipping it, since a failed cheap attempt followed by an expensive one costs strictly more than going straight to the expensive one, and it only pays off when the cheap attempt succeeds often enough
  • Found a defect where approved content was appended without limit to context injected into every prompt, so each approval permanently raised cost and latency for every future call while diluting the instructions that mattered. Fixed by bounding it and by making the size visible to the person creating it
  • Restructured prompts so stable content comes before variable content, which is what makes caching possible at all and costs nothing in quality

Visuals

Two kinds of check: what blocks a merge, and what needs a human
Building an evaluation set from outcomes that already exist, and what must be excluded
Routing work between a cheap and an expensive model, and where it stops paying
Context growth, bounded against unbounded, and what it costs per call