All articles
SPEND CONTROL

AI spend control for agencies: a practical playbook

5 minutes read

AI spend control for agencies: a practical playbook
Summary

Agencies have a harder AI cost problem than most: the spend belongs to clients, not the company. Here is how to attribute it, bill it, and cap it before it eats margin.

Agencies have a harder AI cost problem than almost anyone else, and most do not see it until it hits a margin report.

The reason is simple. In a normal company, AI spend is an internal overhead. In an agency, most of it is done on behalf of clients. The copy your team drafts, the research they run, the analysis they produce, all of it consumes AI on a specific client's work. If you cannot say which client, you cannot bill it, cap it, or defend it. You just absorb it.

This is a playbook for getting AI spend under control when the spend belongs to your clients, not to you.

Why agencies get hit hardest

Three things stack up in an agency that do not stack up elsewhere.

The spend is client work, but it looks like overhead. Without attribution, AI cost sits in one company-wide bucket. Profitable clients and unprofitable ones look identical, so you cannot tell which engagements are quietly losing money.

Usage is spiky and per-project. A single large project can consume more AI in two weeks than the rest of the agency does in a month. A flat company budget cannot see that coming.

Client data raises the stakes. Your people paste client material into AI tools. For an agency, one mishandled dataset can end the relationship. Spend control and data governance are the same problem here, not two.

Step one: attribute before you do anything else

You cannot control what you cannot attribute. Every other step depends on this one.

Tag AI usage to a client and project as it happens, not at month end. Live attribution turns a single unreadable AI bill into a per-client line you can act on. Reconstructing it later from memory is guesswork, and guesswork does not survive a client conversation.

Once usage is attributed, you can finally answer the question that matters: what does AI actually cost us to serve each client?

Step two: choose showback or chargeback

With attribution in place, you have two ways to handle the cost. Most agencies use both, for different clients.

ModelWhat it meansBest for
ShowbackYou track and report AI cost per client internallyRetainers, fixed-fee work, understanding margin
ChargebackYou bill AI cost through to the client directlyTime-and-materials, pass-through, larger projects

Showback is the minimum. Even if you never bill a client for AI directly, you need to know what each one costs to serve, so you can price the next engagement properly.

Chargeback is where attribution pays for itself. If your contracts allow pass-through costs, billing AI back at cost, tagged per client, turns an untracked overhead into a clean line on the invoice.

Step three: set limits per team and per client

Attribution tells you what happened. Limits control what happens next.

  • Per-team limits keep each part of the agency inside its budget. Delivery, creative, and strategy carry different AI loads and should carry different caps.
  • Per-client limits protect margin on fixed-fee work. If a client engagement is scoped for a set amount of AI, the limit should hold at that amount, not warn you after you have blown the scope.

The limits have to be hard. A soft alert on a client project is a margin leak with a notification attached. The point of a cap on fixed-fee work is that it stops.

Step four: govern access before it becomes a problem

Governance is not separate from spend control in an agency. The same controls that manage cost manage risk.

  • Assign models by role, so juniors are not running expensive reasoning models on routine work.
  • Keep audit logs, so you can show a client exactly how their data was handled if they ask.
  • Keep processing inside a compliant boundary. For European clients, EU data residency is often a contractual requirement, not a nice-to-have.

An agency that can show a client its governance wins trust that a competitor guessing at spend cannot. For the underlying budgeting mechanics, see how to budget AI usage across teams.

Step five: stop runaway spend before it starts

Runaway AI spend in an agency almost always comes from one of three places:

  1. Default-to-strongest. Routine work running on the most expensive model. Fixed by assigning models to the task.
  2. The unscoped project. A large engagement with no per-client cap. Fixed by a hard limit tied to scope.
  3. The forgotten workflow. An automation quietly burning tokens on nobody's budget. Fixed by attribution, which makes it visible.

Catch these with a weekly look at spend by client, not a monthly one. On agency margins, a month is long enough for a single unscoped project to erase the profit on three others.

The playbook in one page

  1. Attribute every unit of AI spend to a client and project, live.
  2. Run showback on every client, chargeback where contracts allow.
  3. Set hard per-team and per-client limits that stop at the number.
  4. Assign models by role and keep audit logs for client trust.
  5. Review spend by client weekly to catch runaway projects early.

Where Hebno fits

Hebno was built with agencies in mind. Every major model behind one interface, live spend by team, person, and client, hard budgets that stop at the limit, and usage tagged to the client it was spent on, passed through at cost with no markup. It is showback and chargeback without the spreadsheet.

See where your AI spend is going.

Most teams find the number surprising. Book a 20-minute demo and we will show you yours.