All articles
AI Insights

LLM Cost Regression Testing in CI: A How-To

A prompt tweak, a model upgrade, or a new tool call can quietly double what an AI feature costs to run — with no error, no crash, and no alert, because nothing “broke.” The bill just goes up. LLM cost regression testing closes that gap by treating cost the same way you already treat correctness: something you check automatically, on every change, before it reaches production. Here’s how to set it up.

Why does LLM cost need its own regression tests?

Traditional regression testing catches functional breakage — a feature that used to work and now doesn’t. Cost regressions are different: the feature still works, the output still looks right, but a change upstream (a longer system prompt, a model swap, an extra retrieval step, a retry loop that fires more often than expected) quietly increases tokens per request. Without a dedicated check, that kind of regression is invisible until someone notices the invoice. Tools built specifically for LLM evaluation — Braintrust, Promptfoo, Langfuse and similar platforms — have converged on treating cost and latency as first-class metrics alongside output quality, precisely because teams kept shipping quality-neutral changes that were quietly cost-negative.

What does a cost regression test actually check?

A cost regression test runs a fixed set of representative inputs through your AI pipeline and asserts that token usage, estimated spend, and latency stay within a defined budget — failing the build if any of them drift past a threshold you set. It’s the same shape as a performance regression test in traditional software, just measuring tokens and dollars instead of milliseconds and memory.

Want the numbers before you build?

Get the free Australian AI MVP Cost Guide 2026 — we’ll email it straight to you.

How do you set up cost regression testing in CI?

  1. Build a fixed evaluation set. Pick 20–50 representative inputs that reflect your real traffic mix — short queries, long ones, the edge cases that tend to blow out context. This set needs to be stable so runs are comparable over time; treat it like a versioned fixture, not something regenerated per run.
  2. Record a baseline. Run the eval set against your current production pipeline and log tokens in, tokens out, estimated cost per model’s published pricing, and end-to-end latency for each input. This baseline is what every future run compares against.
  3. Wire the eval into your CI pipeline. Run the same eval set on every pull request that touches prompts, retrieval logic, tool definitions, or model configuration. Platforms built for this — Promptfoo’s CLI-first workflow and Braintrust’s CI integrations are both designed to run as a pipeline step rather than a separate manual process.
  4. Set a drift threshold, not a hard cap. A single flaky retry can spike one run’s token count without meaning anything. Compare against a rolling baseline (e.g. median of the last 5 runs) and fail the build on a sustained percentage increase — commonly somewhere in the 10–20% range — rather than any single-run number.
  5. Break cost down by call, not just by request. One user-facing request often triggers several model calls — a retrieval step, a tool call, a final generation. Track each separately so a regression test can tell you which specific call got more expensive, not just that the total did.
  6. Alert on the pull request, before merge. The entire point is catching the regression before it ships, not in a weekly cost report after a thousand customers have already paid for it. Post the comparison as a PR comment or CI check, the same way you’d surface a failing unit test.
  7. Re-baseline deliberately, not automatically. When a cost increase is intentional — you upgraded to a stronger model because the quality gain justified it — update the baseline explicitly as part of that change, with the reasoning in the commit. An auto-updating baseline defeats the purpose: it would silently absorb every regression as the new normal.

Free: The Australian AI MVP Cost Guide 2026

Honest cost benchmarks, the hidden costs vendors don’t quote, and a 10-line scoping worksheet.

Get the free guide

What tools handle this today?

The LLM observability and eval space has matured specifically around this problem. Braintrust and Promptfoo both publish direct comparisons of their approaches to CI-based evaluation, and the broader observability category — including Langfuse and Helicone — now treats cost tracking as a standard feature rather than an add-on. None of these tools will pick the right threshold or evaluation set for you; that judgment call is still the team’s to make. What they remove is the manual work of running the comparison and wiring it into your existing pipeline.

If you’re not ready to adopt a dedicated platform, the same discipline works with plain scripts: log token counts and estimated cost to a file or table on every CI run, diff against the last known-good baseline, and fail the build past your threshold. The tooling is a convenience, not a prerequisite — the habit of testing cost the same way you test correctness is what actually prevents the regression from reaching customers, and it pairs directly with the retrieval-tuning work in our guide to RAG chunking and retrieval tuning, since a bad chunking change is one of the more common silent cost regressions in a retrieval pipeline.

How does this fit with functional evals and red-teaming?

Cost regression testing isn’t a replacement for output-quality evals or security testing — it runs alongside them. A pipeline that’s cheap but wrong, or cheap but exploitable, hasn’t solved anything. If you haven’t set up either yet, our guides to building an AI product cost model and red-teaming an LLM app cover the other two legs of the same production-readiness stool: know what a request should cost before you optimise it, and know what a request could be tricked into doing before you scale it.

Frequently asked questions

What is LLM cost regression testing?

It’s an automated check, run in CI on every relevant change, that measures token usage, estimated spend and latency for a fixed set of representative inputs and fails the build if they drift past a set threshold — the same discipline as a performance regression test, applied to cost instead of speed.

How big should the evaluation set be?

20–50 representative inputs is a practical starting range for most products — enough to cover your real traffic mix (short queries, long ones, known edge cases) without making every CI run slow or expensive to execute.

Should a cost regression test use a hard cost cap or a drift threshold?

A drift threshold compared against a rolling baseline works better in practice than a hard cap, because a single flaky retry can spike one run’s numbers without indicating a real regression. Failing the build on a sustained percentage increase against a rolling baseline (commonly in the 10–20% range) avoids both false alarms and missed regressions.

Do I need a dedicated tool like Braintrust or Promptfoo to do this?

No — a plain script that logs token counts and cost to a file and diffs against the last known-good baseline achieves the same result. Dedicated platforms remove manual wiring and add reporting, but the underlying discipline works without them.

Neomeric, a Melbourne-based AI product and consulting company — and the team behind NeoMind, Australia’s onshore AI teammates platform — builds this kind of production discipline into every AI product it ships.

Sources

Building something? Get a straight answer on cost.

Neomeric is a Melbourne AI product studio — 7+ products shipped, including our own. Start with a free 15-minute scoping call, or a 2-week Build Sprint at A$6,900 fixed, fully credited toward your pilot.

Book a free scoping callDownload the cost guide

Disclaimer: This article is general information only, current at the time of writing, and is not legal, financial or professional advice. Regulatory obligations, pricing and market figures change and vary by circumstance — seek advice specific to your situation before acting. Statistics cited are drawn from the third-party sources linked in this article; Neomeric is not responsible for third-party content.

AI Insights AI Strategy Australian AI
PDF · Free

Get the Australian AI MVP Cost Guide 2026.

What an AI MVP really costs in Australia in 2026 — line-item budgets, the traps that blow them out, and how to scope a build that pays for itself.

N
Neomeric Team

We build the AI products others can’t. Melbourne, Australia. Work with us →