Subscribe to Newsletter

Module 5: Token Spend and Economics

Learn to govern spend before you can measure impact and find who drives the bill. Then keep it flat with controls that act before a request, not after.

Start module Module 5 of 5 · 4 lessons

5.1

How to set an operating model for AI spend when you cannot yet measure impact

Uber is the one large company whose budget failure played out in public, with its executives on the record. The sequence:

  1. Uber encouraged staff to use AI as much as possible and ranked teams on an internal leaderboard by total AI tool usage.
  2. In April 2026 the CTO disclosed that the company spent its entire annual AI coding-tool budget in four months.
  3. In June, Uber set a monthly $1,500 cap per employee and per agentic coding tool, visible on a dashboard and exceedable with permission.
  4. Uber’s president and COO admits the link between spend and shipped features “is not there yet.”
THE TARGETTeams ranked on aleaderboard by totalAI tool usageAPRIL 2026Annual AI coding-toolbudget spent infour monthsJUNE 2026$1,500 monthly cap peremployee, per agenticcoding tool; exceedablewith permissionON THE LINK BETWEEN SPEND AND SHIPPED FEATURES“That link is not there yet.” Andrew Macdonald, Uber president and COO
Scroll sideways →Sources: Fortune, TechCrunch.

Uber’s leaderboard is the mechanism the lesson on adoption versus impact warned about: a usage metric turned into a target.

Gartner’s five-part operating model for agent use

An operating model decides who uses agents for what, which model handles each task, and who reviews the token bill. Most organizations still lack the maturity and frameworks to measure cost against business impact. Gartner’s answer is a disciplined operating model for AI usage, in five parts:

  1. Establish a use-case-driven decision framework. Define when teams use agents and at what autonomy. Classify tasks as developer-led, developer-with-agent, or fully agent-led.
  2. Align model selection with task complexity. Route simple, high-frequency tasks to smaller models. Reserve frontier models for complex, high-value work.
  3. Mandate context engineering practices. Train developers to include only relevant information, summarize content where possible, and remove unnecessary data.
  4. Implement governance and cost controls. Set token thresholds, escalation policies, and automated monitoring.
  5. Embed token usage reviews into development cycles. Review high-token workflows in sprint retrospectives.
TRY THIS TODAY

Goal: Find whether anyone measures each team’s spend against an outcome, and whether last month’s spend was accidents or sustained work.

  1. Export last month’s token spend by developer from your vendor consoles or gateway.
  2. Paste the export into your coding agent with this prompt:
    Prompt
    Here is one month of AI token spend by developer, with team.
    
    1. Compute total spend, spend per active developer, and the share of spend
       from the top 10% of developers.
    2. For each developer in the top 10%, mark whether their spend looks like
       an accident (one or two days far above their usual) or sustained
       (elevated most working days). Use the daily data if present.
    3. For each team, state whether spend is currently compared against any
       outcome (merged PRs, cycle time, shipped items). Answer "unknown" if
       the data does not say.
    Return three tables. Do not add recommendations.
  3. Mark each team that answered “unknown”. Uber sat there before the cap.
  4. Give each accident an alarm.
  5. Give each sustained spender a named owner and a budget line.

Expected result: spend per active developer, the top 10% share, each outlier marked accident or sustained, and each team marked by whether anyone measures it.

5.2

How to set budgets and measure what you bought

A runaway script and a six-week migration can both blow past an engineer’s usual spend. The script needs an alarm. The migration needs an owner and a budget. If you give those two cases the same control, you will get the worst of both: accidents that run for hours and real work that waits for approval.

Separate accidents from sustained spend

A daily guard interrupts an unusual burst and asks a human to look at it. A monthly limit governs sustained usage. Either way, show engineers what they have spent across every tool before a limit turns into a surprise.

Databricks’ budget design came out of a failure. Under a single monthly cap, 500 to 1,000 engineers hit the limit every month and filed tickets to raise it, and every raise was permanent.

The replacement lets engineers clear daily increments themselves. Larger monthly allowances need a manager’s approval and expire with the project. Raising the monthly tier raises the daily increment with it. The pattern keeps ordinary work moving while making unusual spend deliberate. Uber’s $1,500 per employee per tool, with a dashboard and an exception process, is the blunter cousin of the same idea.

DecisionDaily guardMonthly limit
PurposeCatch an unexpected burst.Approve a sustained level of spend.
When reachedShow the recent spend. Ask a human to inspect it and resume intentionally.Ask the manager to approve a named project and an expiry date.
ScopeAggregate usage across tools where metering permits.Include every tool in the team’s overall envelope.
ResetReturn to the base daily allowance on a published schedule.Revert temporary increases when the project ends.
TWO BUDGETS, TWO KINDS OF WASTE DAILY · RUNAWAY GUARD 90% LIMIT self-serve · unlimited acks · resets nightly You’ve spent $X today, $Y headroom. One click raises the limit one increment. This is intentional effective_limit = min(spent + one daily increment, monthly max) Raise the monthly tier and the daily increment scales with it, so the runaway guard stays calibrated. MONTHLY · EXTRAORDINARY SPEND Base ~2x ~5x Effectively unlimited each step: manager approval · project-scoped · 1 / 3 / 6 months, then reverts BELOW SUSPENSION Downshift to cheaper model Suspend (last resort)
Scroll sideways →Source: Databricks.

Make the exception path part of the policy

Do not wait for the first production incident to decide who can override a limit. Define the emergency path now, log every use of it, and review it afterward. A security incident may justify an immediate suspension. Ordinary overspend deserves an explanation and a way back.

Downshifting means offering a cheaper model when a limit is reached. Use it wherever that model can still do the work safely, and tell the engineer what changed. Silently swapping model capability is the fastest way to turn a budget surprise into a debugging surprise.

budget-policy.md
Daily guard: [threshold and increment based on observed usage]
Acknowledgment: [where a human inspects spend and resumes]
Monthly base: [allowance] | temporary tiers: [allowances]
Override owner: [manager] | project: [name] | expires: [date]
Emergency path: [on-call owner and audit trail]
Downshift: [eligible tasks and available model]
Review: weekly exceptions; monthly policy calibration

Connect spend to accepted work

A smaller bill is easy to celebrate. It is harder to notice that the team shipped less, or spent Friday repairing generated code. Your measurement has to catch both, and minware’s framing is the right one: delivery per dollar is the number the board cares about.

Define “accepted” before you count

An accepted task satisfies the intended change and passes your review bar. A pull request that exists is not the same as a task that is done. For a documentation task, acceptance might mean a factual review. For a migration, it might mean validation against production-like data. Write the bar down before you count anything.

cost-per-accepted-task
Model-only cost per accepted task
= model spend across all attempts / accepted tasks

Fully loaded cost per accepted task
= (model spend + allocated tool fees
   + repair hours × loaded hourly rate) / accepted tasks

Use the same accounting boundary for every comparison, and allocate shared fees the same way each time. If no tasks pass, report the spend and zero accepted work. Do not let a dashboard show a reassuring zero-dollar unit cost.

Build a weekly view a team can act on

Total spend and accepted tasksShows the bill beside what was delivered.Check whether workload or task difficulty changed.
Cost per accepted task by task classPrevents easy tasks from hiding expensive failures.Inspect failed attempts and repair work in that class.
Median and p90 session costThe median is the middle session. The p90 marks the cost below which 90% of sessions fall.Read traces from the costly tail. Averages can conceal it.
Cycle time and rework rateChecks whether cheaper output creates slower delivery.Pause an efficiency change if it moves work onto reviewers.
Unattributed spendShows what you cannot connect to a team or task.Fix the missing identity or task mapping before declaring ROI.

Higher delivery next to higher AI spend is a clue, not proof. Teams differ in workload and staffing, so compare similar work over time and run a controlled pilot when you can. Faros makes the same point about leading indicators: the share of traffic on frontier models should fall as routing matures.

Never turn this view into a ranking of engineers by token usage or raw pull-request count. Uber did, and it ended with a cap.

Stress-test the budget

A forecast should answer “what if” before it pretends to answer “how much.” Nobody, including the provider, knows what a request costs before it returns, which is why the 2026 FinOps Framework asks for a range with stated assumptions rather than a single number.

Start with your active engineers and your observed monthly spend, then change one assumption at a time. The calculator keeps headcount fixed and grows the active share only until everyone is active.

The baseline is consumption spend only. Add fixed subscriptions and infrastructure separately. Set the shock inputs to zero to see today’s envelope, and turn every savings lever off before you estimate an improvement.

For calibration, one major provider’s enterprise data puts the average at $150 to $250 per developer per month. Uber saw $150 to $2,000 before its cap, and Gartner warns of $2,000 to $5,000, with extremes at $20,000.

Spend envelope calculator
Shocks
Levers
Baseline
$0
$0 / year
After shock
$0
$0 / year
×1.0 vs baseline
After levers
$0
$0 / year
Annual envelope
Baseline$0
After shock$0
After levers$0

This is a sensitivity model, not a savings promise. Each lever reduces the remaining spend. Routing and cheaper defaults may overlap: switch one off unless you have measured separate gains. Fixed fees and human time are excluded. Adoption growth is capped at 100% of engineers.

5.3

How to find which engineers drive your token spend, and baseline it before rollout

A few engineers running agents in loops can account for most of your token bill, so the average describes nobody.

Jellyfish names four variables that set one engineer’s spend. Headcount is not one of them:

  1. Model selection. Frontier reasoning models cost an order of magnitude more per token than small, fast models.
  2. Context size. Large system prompts and wide context windows push more tokens into every request.
  3. Retrieval strategy. Weak retrieval pulls in far more files than the task needs, and the model reads all of them.
  4. Agent autonomy. One task can trigger a long run of planning, editing, testing, and retrying. Each step bills as another call.

Each variable multiplies the others. A developer picks the frontier model, works with wide retrieval, and runs agents unattended. That developer can spend a hundred times what a colleague spends on the same tool.

Model selectionfrontier vs small×Context sizelarge prompts,wide windows×Retrievalwide vs narrow×Agent autonomyunattended loopsvs single callsUp to 100x between two engineers on the same tool
Scroll sideways →Source: Jellyfish, AI Token Usage Monitoring.

The distribution that results has a long right tail. The tail pulls the average up, and the average still understates the tail.

Two numbers replace the mean. The first is cost per developer per active day, which normalizes for how often people use the tool and which vendors report themselves. The second is the P90, the cost below which 90% of sessions fall. It exposes the expensive tail that an average hides, as our costs guide’s lesson on the weekly view shows.

MEDIANabout $52MEANpulled up by the tailP90close to $691Top 10% of developersMONTHLY TOKEN SPEND PER DEVELOPERNOT TO SCALE
Scroll sideways →Source: Jellyfish, AI Token Usage Monitoring.

1. How to baseline token spend per developer before rollout

A baseline is the spend shape of a small group, measured before rollout. It has three parts: the average per active day, the monthly range, and the P90.

Anthropic publishes that shape across its enterprise customers. The average is about $13 per developer per active day and $150 to $250 per developer per month. Costs stay below $30 per active day for 90% of users.

Start with a small group and use the tracking tools to set a baseline before wider rollout.

Where you can see, cap, and attribute spend depends on the setup. Only one route works everywhere. OpenTelemetry export works on every setup. It is the only option that streams per-user token and cost metrics into your own observability stack in near real time. The in-product cost figures are estimates, not an invoice.

Your setupSee spendCap spendPer-user reporting
Claude for Teams or EnterpriseSpend report in org analyticsSpend limits in admin settingsSpend report CSV; Enterprise Analytics API on Enterprise
Claude Console (API)Console usage pageWorkspace spend limitsConsole dashboard, Claude Code Analytics API
Amazon Bedrock, Google Cloud’s Agent Platform, or Microsoft FoundryYour cloud billing consoleYour cloud’s budget controlsOpenTelemetry or an LLM gateway

The DX pricing guide puts heavy users on subscription tiers rather than raw API rates.

2. Why your top spenders cost 13 times the median

The P90-to-median ratio measures how far the heavy users sit from the typical one. The wider the ratio, the more an org-level total hides.

Across the customers in that dataset, the median developer spends about $52 a month on tokens. The 90th percentile spends close to $691, roughly thirteen times more. In tokens, that is about 51 million a month for the median developer and roughly 380 million at the 90th percentile.

A small group of power users accounts for most of that total, and an org-level number hides every one of them.

STEP 1Pilot groupA small group, measuredbefore wider rolloutSTEP 2Cost per developerper active dayReported as medianand P90STEP 3A per-team alert atthat team's P90
Scroll sideways →Sources: Anthropic, Manage costs effectively, Jellyfish, AI Token Usage Monitoring.

A worked example shows what the tail does to a budget:

DevelopersMonthly token spend
200-developer org, everyone at the median200 at $52about $10,000
Same org, twenty developers move to the P90180 at $52, 20 at $691past $23,000

The org hired nobody and changed no roadmap.

The tail also buys less per token. By decile:

Bottom decileMedianTop decile
Merged PRs per engineer per week0.772.15
Tokens per PR7.5 million69 million
Cost per merged PR$0.28 (lightest users)$89.32 (heaviest users)

That is ten times the tokens for a bit over twice the work.

PRS PER ENGINEER PER WEEK 0.77 0.85 0.92 1.04 1.00 1.00 1.08 1.31 1.54 2.15 TOKENS PER PR 74K 691K 1.9M 3.8M 7.5M 12M 17M 24M 35M 69M D1 D2 D3 D4 D5 D6 D7 D8 D9 D10 ← LOW TOKEN USEHIGH TOKEN USE →
Scroll sideways →Source: Jellyfish, AI Token Usage Monitoring.

A high spender can be a productive one. Productive developers tend to use more of everything.

3. Seven metrics for a token usage dashboard

A usage dashboard shows where the tokens go, and pairs each number with the sign that something is wrong. A working dashboard has seven metrics:

Metric
Metric 1Total token consumption

What it tells you. This is the number finance works from.

Warning sign. It climbs faster than headcount and roadmap scope.

Metric 2Usage by tool

What it tells you. It shows how the budget splits across assistants and agents.

Warning sign. One tool holds a large share while few developers use it.

Metric 3Usage by team and user

What it tells you. It shows where consumption concentrates.

Warning sign. A small group carries a share of spend well above its share of the work.

Metric 4Input, output, and cache mix

What it tells you. It shows how well tools handle context.

Warning sign. Input climbs faster than output, or cache reads sit near zero.

Metric 5LLM call volume

What it tells you. It shows traffic, separate from token weight.

Warning sign. Calls rise faster than tokens, which points to retries and agent loops.

Metric 6Response latency

What it tells you. It shows whether developers keep the tool on.

Warning sign. A p95 slow enough that developers turn features off.

Metric 7Spend by provider and model

What it tells you. It gives you renewal leverage.

Warning sign. Premium models run work that a cheaper model handles well.

To tie the numbers to output, add four team-level measures:

  1. Tokens per PR only means something against a team’s own history. A typo fix and a 2,000-line refactor both count as one PR.
  2. PR cycle time shows whether AI-assisted work clears review faster.
  3. Code review activity shows the quality effects that raw throughput hides.
  4. Roadmap work against keep-the-lights-on work tells the two apart. Two pods that burn 100 million tokens each, one on maintenance and one on feature work, produce identical invoices and very different quarters.

Start with visibility before optimization.

The usual triggers are a workflow change, a model change, a new user, and a retry loop. Alerts work on percentage change in daily or weekly consumption, and thresholds per team catch more than a single org-wide threshold. The quiet case is a system prompt change that doubles input tokens and goes unnoticed for weeks.

4. Why you cannot explain a token bill without a baseline

A baseline is the measurement taken before the tool spreads. Without one, later spend has nothing to compare against.

Without a baseline today, you cannot measure impact, optimize spend, or defend the investment when leadership asks whether AI moves the needle. You will have a bigger bill and no answer. If your bills look manageable today, stress test what they look like in Q4.

A bill can also change while usage stays flat. Pylon’s Anthropic spend was on track to jump from $400K to $1.4M a year. Usage did not grow. The company crossed a seat threshold that triggered a tier change.

For the business case, three inputs finance will accept:

  1. Fully loaded cost per engineer. Seat plus token spend plus governance and infrastructure plus onboarding time, at the finance team’s loaded developer cost.
  2. Validated time savings. Developer-reported hours saved per week, checked against task tracking. Apply a conservative utilization factor.
  3. Quality-adjusted productivity. Track rework rates, review iteration counts, and incident rates for AI-touched code beside velocity metrics.

The ranges for the first input:

InputRange
Onboarding time, first quarter4 to 8 hours per engineer
Onboarding time, each later quarter1 to 2 hours per quarter
Loaded developer costtypically $75 to $100 per hour
TRY THIS TODAY

Goal: Produce your own median, P90, and top-10% share from one month of data, and set one alert.

  1. Export one month of token spend per developer per day from your vendor console, gateway, or OpenTelemetry sink.
  2. Paste it into your coding agent with this prompt:
    Prompt
    Here is one month of AI token spend per developer per day.
    
    Compute, per team and for the org:
    - cost per developer per active day (days with any spend)
    - the median and the 90th percentile of monthly spend per developer
    - the share of total spend from the top 10% of developers
    Report the P90-to-median ratio. Do not report the mean on its own.
    Then list the developers above the P90 with their daily pattern:
    sustained (elevated most days) or spike (one to three days).
  3. Put your P90-to-median ratio beside the industry ratio of thirteen.
  4. Put your cost per active day beside the enterprise figures of $13 and $30.
  5. Set a per-team alert at the team’s own P90, on percentage change week over week.
  6. Ask the top spender’s team three questions: more shipped, faster review, code that held up.

Expected result: a per-team median and P90 against the industry shape, and the top-10% share of spend. Add one alert per team at its P90 and three answers for the heaviest team.

The Code: Your daily unfair advantage in software engineering.

Join 350,000+ software engineers, tech leads, and CTOs who start their morning with The Code.

Subscribe to Newsletter
5.4

How to keep your token spend flat as usage grows, with one gateway

Token usage in your organization will grow every quarter. Spend does not have to grow with it. Coinbase cut its AI spend nearly in half while usage kept rising. Shopify keeps its bill flat with a $250-a-day question, and PostHog does it with a cache that most teams get backwards.

This lesson covers what those three did, and the gateway that made each of it possible.

1. Why one gateway should sit between every tool and every model vendor

Each AI request from an engineer’s tools goes to some model at some price. If each tool talks to each vendor directly, there is nowhere to set a policy once.

If every request passes through one proxy, that proxy can do four things for the whole organization at once:

  1. It sets the default. A request uses the model the tool ships with, and the gateway decides what that is.
  2. It routes each step. A planning step may need a frontier model, and the execution steps that follow usually do not.
  3. It caches. The cache can serve requests that share a prefix, such as the system prompt and tool definitions, at a tenth of the price. The gateway can enforce the prompt order that makes that happen.
  4. It sees every request. Each one carries a team, a project, and a person, and the dashboard, the alert, and the chargeback all come from this one stream.

The FinOps Foundation’s token economics paper, a working-group publication with contributors from Shutterstock, AWS, Wayfair, Superhuman, and Align Technology, recommends the gateway. A centralized gateway enforces model selection policy, injects cost metadata, and exposes routing configuration to application teams. The paper calls it the single architectural decision most likely to produce sustained, org-wide savings.

Watch: How Coinbase cut AI inference costs by half while tokens went up

2. How Coinbase cuts token spend

A change to a default, a route, or a cache moves the whole bill at once. It acts on every request before the tool sends it.

Brian Armstrong, CEO of Coinbase, laid out how Coinbase keeps AI spend flat while token usage grows exponentially: “Not with friction and spend alerts. With better defaults, routing, and caching.”

His five practices, under his headings:

  1. Better Defaults (not Usage Caps). Coinbase defaults to open-weight models through its gateway. 91% of its employees never hit their usage caps, so instead of lowering caps and driving up alerts, it moved to cheaper defaults.
  2. Better Routing. A frontier model for planning, not for execution. Humans should not choose models. AI can automate that task.
  3. Better Caching. Every request is cache aware. Coinbase’s cache hit rate went from 5% to 60% in LibreChat once it implemented caching properly.
  4. Keep Context Lean. Fresh sessions per task, narrow file scope, unused tools disconnected.
  5. Better Visibility. Engineers may use any model at any volume. Coinbase makes usage visible, and the more an engineer spends on AI, the more impact it expects.

The practices cut Coinbase’s AI spend nearly in half while its token usage continues to grow.

BEFORE THE REQUESTEvery toolEvery modelONE GATEWAYDEFAULTOpen-weight models;91% never hit their capsROUTEFrontier for planning,not for executionCACHEEvery request cache aware;hit rate 5% → 60%SEEUsage visible; more spendmeans more impact expectedFifth practice: keep context leanAFTER THE SPENDCapsAlertsChargebackCoinbase: AI spend nearly halved while token usage keeps growing
Scroll sideways →Source: Brian Armstrong, CEO of Coinbase, on X.

Box’s CEO added the condition that none of it works without a deep, concrete understanding of the underlying work.

3. How Shopify cuts token spend

An alert can block a spend, or it can open a conversation about it. Shopify chose the conversation.

Shopify’s VP and head of engineering applies one rule in Bessemer’s account of its playbook: standardize infrastructure, not tools. A central proxy routes every request from every tool, Shopify buys tokens in bulk, and usage shows by team, by project, by person.

Shopify gets an alert when someone spends more than $250 in tokens in a day. The head of engineering then investigates what they build. He often finds ambitious and worthwhile experiments, such as attempts to refactor large parts of Shopify’s mobile codebase.

The circuit breaker, from VentureBeat’s interview, works the same way. When a user has a model running for around ten hours and it consumes many tokens, Shopify pings them with one question: “Did you mean to spend this?” Some say yes. Others did not know it still ran in the background and stop it.

The gateway also does what a cap cannot: it swaps models without anyone changing a tool. When Claude Fable 5 shut down, Shopify’s engineers did not panic. The proxy shifted them to Claude Opus or GPT 5.5 automatically.

Shopify also distills frontier models into small ones for specific subtasks, through a pipeline an engineer runs in about a day.

THE PROXY ACTSTHE ALERT ASKSEvery toolCentral proxyRoute to a modelSwap automaticallyon outageSend to a distilledsmall modelUsage by team, project, personAlert: over $250 in a day,or a model running 10 hours“Did you mean to spend this?”Yes: keep goingNo: stop it
Scroll sideways →Sources: Bessemer, Inside Shopify’s AI-first engineering playbook, VentureBeat.

The head of engineering’s own caveat: spend rises non-trivially against his humble estimate of a 20% productivity gain.

4. How PostHog cuts token spend

Token spend hides in three places. PostHog tracks all three: flat-rate subscriptions, internal automations, and customer-facing AI features. Teams almost always have a gap in at least one of them, and the monthly bill from Anthropic will not show it.

The cache is the easiest lever to get backwards. A fresh session that cut accumulated input by 89% raised cost, because every new call had to rebuild the whole cache. Five rules follow:

  1. Caches are per model, so a mid-session fallback to another model doubles the charge at once.
  2. The cache key is the literal bytes of the prompt in render order. Put stable tokens first and volatile content last.
  3. Subagents do not share the parent’s prefix, so each one pays the full cache-write price.
  4. A write costs more than normal input and a read a fraction of it. One read earns back the five-minute write.
  5. Check cache_read_input_tokens in the response. If it is zero across requests that should share a prefix, something silently invalidates the cache.

Model choice follows the same test: measure quality, speed, and cost together, not the list price.

PostHog’s test, Sonnet 5 against Opus 4.5
Advertised2.5x cheaper
Measured20% cheaper
ReasonSonnet spends 1.7x the turns and tokens

Cheap models take one-shot jobs such as titles and classifiers. Frontier models take work where the reasoning is the product. Per-request routing across models is the current meta, with some companies saving 60 to 75% on workloads.

The last rule is when not to optimize. Look at per-run cost to find waste and at run count to decide if a change is worth it. PostHog’s install wizard runs hundreds of times a day at about $7 each. A prototype does not qualify, because teams rewrite prototypes.

5. Where different companies have placed their token controls

Each org prefers a different place to put their token controls. This is where some of the companies we mentioned placed their controls:

  • Gartner: thresholds and escalation policies, because token discipline will not emerge through developer choice alone.
  • Armstrong: defaults rather than friction and spend alerts, because most engineers never hit their caps and the defaults do the work.
  • Shopify: an alert at $250 a day that opens a conversation, and a circuit breaker at ten hours.
  • FinOps: a budget set above a measured baseline with two alerts, given in the budget rule.
ARMSTRONGDefaults, routing,and cachingSHOPIFYA $250 question;a ten-hour breakerFINOPSBaseline, with 80%and 100% alertsGARTNERThresholds andescalation policiesBEFORE THE REQUESTAFTER THE SPEND
Scroll sideways →Sources: Armstrong, FinOps Foundation, VentureBeat on Shopify, Gartner.
TRY THIS TODAY

Goal: Make three gateway changes in one week and measure each.

  1. Default. Define the quality bar for one team’s most common task type.
  2. Run the top three candidate models against twenty recent tasks.
  3. Set the gateway default to the cheapest model that passes.
  4. Record the before-and-after cost per task.
  5. Alert. Add a per-person daily alert at Shopify’s $250, worded as a question: “Did you mean to spend this?”
  6. Route it to the person and their manager.
  7. Record how many fire in the week and how many were intended.
  8. Cache. Pull the last week of requests for the five highest-volume workflows.
  9. Paste the cache fields into your coding agent with this prompt:
    Prompt
    Here are request logs for five workflows with cache_read_input_tokens,
    cache_creation_input_tokens, input_tokens, and model per request.
    
    For each workflow report: cache hit rate (cache reads / total input),
    the share of requests with zero cache reads, and whether the model
    changed mid-session. Flag any workflow under 30% hit rate or with a
    mid-session model change. Do not add recommendations.
  10. For each flagged workflow, move stable content to the front of the prompt, per the second cache rule.
  11. Move volatile content after the last cache breakpoint.
  12. Re-measure the hit rate.

Expected result: one default changed with a cost-per-task delta, and one alert that asks instead of blocks, with a count of intended versus unintended spend. Add cache hit rates for five workflows with the fixes applied.

END OF MODULE 5

By this point you should have:

  • Each of your teams marked by whether anyone measures its spend against an outcome, and last month’s outliers sorted into accidents and sustained spend.
  • Gartner’s five-part operating model to score against.
  • A daily guard and a monthly limit with a written exception path, and a cost per accepted task.
  • A weekly view that pairs spend with delivery, and a stress-tested spend envelope.
  • Your own median, P90, and top-10% share, with a per-team alert at the P90.
  • One gateway default changed, one alert that asks instead of blocks, and cache hit rates measured and fixed on five workflows.

This guide is one argument. Adoption is not impact, and the numbers that show the difference are the ones to report. The bottleneck moved from writing code to verifying it, so review, specs, and enforcement are where the investment goes.

Agents need a bounded blast radius, not a better prompt. Each of those controls has a token bill that no longer tracks headcount. The operating model for spend is the last thing a CTO puts in place and the first thing a board will ask about.

You should now be able to decide where your organization sits on each of those questions. You should know what to change next, and what it costs.