← All workflows

Workflow · October 6, 2026

Build a Structural AI Cost Model: Audit Your LLM Spend as Token Prices Deflate

✓ TestedFinanceFor Finance
Time saved4-6 hours per quarterly reforecast

The task

You're the FP&A lead or controller who owns the AI/LLM line in the IT budget. Token prices are falling fast but your invoice keeps climbing, and someone on the exec team is going to ask why before close. This workflow builds a defensible cost model you can bring to the CFO and to vendor renewals — pasted usage data in, structural variance analysis out.

Before AI

Today this is a half-day of spreadsheet archaeology: export usage CSVs from each AI vendor portal, hand-build a pivot by model and team, pull last year's unit prices from old invoices, and manually compute price-vs-volume variance. Then you draft a narrative memo for the business review. Call it 4-6 hours, more if you have multiple vendors on different billing cycles. The problem isn't the math — it's that per-token prices moved so much (one analysis of 2.4 billion enterprise API calls shows the blended cost of AI dropped 67% year over year, from $18.40 to $6.07 per million tokens between Q1 2025 and Q1 2026) that your prior-period baseline is already stale by the time you finish. Deeper context in this breakdown of why bills keep rising anyway.

The workflow

1. Decompose the bill into price, volume, and mix

Paste your usage extract into the first prompt. It will separate the structural drivers so you can see what's actually moving the number.

Prompt
You are a senior FP&A analyst building a variance bridge for LLM/AI spend. You will receive a usage dataset with monthly rows per model, per team, with tokens consumed and dollars billed.

Do the following:
1. Compute implied unit price per million tokens for each row (spend / tokens * 1,000,000).
2. Build a Q-over-Q bridge for total spend decomposed into: (a) price effect — holding volume and mix constant, how much did unit prices change the bill, (b) volume effect — total token growth at prior-period prices, (c) mix effect — shift between cheaper and more expensive models.
3. Flag any model where unit price INCREASED quarter over quarter — that is a renegotiation signal.
4. Call out any team whose token consumption grew more than 3x — that is almost certainly an agentic workflow that was not in the original budget.

Output a clean markdown table for the bridge, then a short bulleted "what to investigate" list. Use the data below.
Sample input
vendor,model,team,month,tokens_millions,spend_usd
Vendor-A,frontier-large,Revenue Ops,2026-04,42.1,1263.00
Vendor-A,frontier-large,Revenue Ops,2026-05,51.3,1488.00
Vendor-A,frontier-large,Revenue Ops,2026-06,58.7,1644.00
Vendor-A,frontier-large,Revenue Ops,2026-07,61.2,1591.00
Vendor-A,frontier-large,Revenue Ops,2026-08,88.4,2210.00
Vendor-A,frontier-large,Revenue Ops,2026-09,214.6,5151.00
Vendor-A,frontier-mini,Revenue Ops,2026-04,8.2,41.00
Vendor-A,frontier-mini,Revenue Ops,2026-09,12.4,56.00
Vendor-B,reasoning-pro,Engineering,2026-04,19.8,1584.00
Vendor-B,reasoning-pro,Engineering,2026-05,24.1,1808.00
Vendor-B,reasoning-pro,Engineering,2026-06,33.6,2419.00
Vendor-B,reasoning-pro,Engineering,2026-07,47.2,3304.00
Vendor-B,reasoning-pro,Engineering,2026-08,62.9,4089.00
Vendor-B,reasoning-pro,Engineering,2026-09,81.5,5134.00
Vendor-B,reasoning-mini,Engineering,2026-04,110.0,330.00
Vendor-B,reasoning-mini,Engineering,2026-09,180.5,487.00
Vendor-C,open-weight-70b,Data Science,2026-04,220.0,264.00
Vendor-C,open-weight-70b,Data Science,2026-09,410.0,451.00
Vendor-A,frontier-large,Customer Support,2026-04,15.2,456.00
Vendor-A,frontier-large,Customer Support,2026-09,19.1,497.00
Vendor-D,legal-domain,Legal,2026-04,2.1,420.00
Vendor-D,legal-domain,Legal,2026-09,2.4,504.00

2. Reforecast the next two quarters against the deflation curve

Now take the bridge from Step 1 and project forward using a conservative structural assumption. We're not predicting the future — we're pressure-testing the budget line against a known trend. Worth remembering why this matters: Menlo found 74% of startups and 49% of enterprises run majority-inference workloads, and Deloitte estimates inference will be roughly two-thirds of AI compute in 2026, up from one-third in 2023, so the inference line is the one to model carefully.

Prompt
Using the bridge and investigation list you just produced, build a forward-looking cost model for Q4 2026 and Q1 2027. Apply these assumptions explicitly and show your work:

- Frontier model unit prices: assume -15% per quarter (conservative; recent reporting suggests steeper structural decline).
- Mini / small model unit prices: assume -10% per quarter.
- Open-weight / self-hosted: assume flat (compute, not license, is the floor).
- Volume: hold each team's September run rate flat UNLESS Step 1 flagged a >3x growth team — for those, model a 1.5x further increase in Q4 then flat in Q1, and label it "agentic workflow risk".

Produce three scenarios: Base (as above), Downside (volume continues growing 50% QoQ for flagged teams), and Upside (one frontier workload migrates to a mini model, cutting that team's spend 60%).

Give me a table with quarterly totals by scenario, then a two-sentence narrative I can paste into the board deck.

3. Draft the vendor renegotiation ask

Use the model to turn the analysis into a specific, defensible counter for your next renewal conversation. Vendors expect finance to know the market now; vague asks get ignored.

Prompt
Based on the variance bridge, the forward model, and any flagged unit-price INCREASES from Step 1, draft a one-page renegotiation brief addressed to our account manager at the vendor with the largest spend concentration.

Structure:
- Opening line stating total spend with them YTD and our projected FY run rate.
- Three bullet points of specific observations (use the numbers from Steps 1 and 2).
- A specific ask: target blended unit price per million tokens for the next 12 months, plus a volume commitment we are willing to make in exchange.
- A walk-away alternative that references migrating workloads to a mini model or open-weight equivalent, with the dollar impact quantified.

Keep it under 300 words. Professional, not combative. No emojis, no filler.

Gotchas

  • Garbage in, garbage out. If your usage export lumps "input tokens" and "output tokens" together, your unit-price math understates the real cost of reasoning-heavy models where output dominates. Pull the detailed cut if you can.
  • Mix shift hides the real story. A team can show "flat spend" while quietly moving from mini to frontier models — the price effect looks neutral but you've taken on a quality/cost commitment that compounds. Step 1's mix line catches this; don't skip reading it.
  • The deflation assumption is not a forecast. The AI subsidy is ending, and waste is about to become a line item — see the full argument from Speakeasy. Treat the -15%/-10% as a planning convention, not a promise. Build the Downside scenario and show it to the CFO.
  • Agentic workflows break the model. A single agentic workflow that looks like "one task" might consume fifty times more tokens than a chat turn. If a team flagged in Step 1 just shipped an agent, your Q4 Base case is already low — 73% of enterprises reported in 2026 that their actual AI costs exceeded original projections, and that is the pattern you're trying not to repeat.
  • Don't send Step 3's draft without a human pass. Vendor relationships have context the model doesn't see — pending contract clauses, SLA credits, roadmap commitments. Treat the output as a first draft.

Time saved

A typical quarterly LLM spend review — pulling usage, building the bridge, writing the memo, prepping the renewal ask — runs 4-6 hours for one analyst. This workflow compresses the mechanical analysis into roughly 30 minutes of prompt-and-review, leaving the saved time for the judgment calls (which teams to confront, which vendor to pressure first). Estimate, not a measurement; your mileage depends on how clean your usage export is.

What we got when we ran it

We executed every prompt above, in order, against the sample input (model: claude-sonnet-4-6). Outputs below are real and unedited apart from length truncation.

Step 1 output

# LLM/AI Spend Variance Bridge — Q2 2026 → Q3 2026 --- ## Step 1 — Implied Unit Price ($/M tokens) by Row > Q2 = months 04–06 aggregated; Q3 = months 07–09 aggregated | Vendor | Model | Team | Q2 Tokens (M) | Q2 Spend ($) | Q2 $/M | Q3 Tokens (M) | Q3 Spend ($) | Q3 $/M | Price Δ | |---|---|---|---:|---:|---:|---:|---:|---:|---:| | Vendor-A | frontier-large | Revenue Ops | 152.1 | 4,395.00 | 28.90 | 364.2 | 8,952.00 | 24.58 | ▼ −4.32 | | Vendor-A | frontier-mini | Revenue Ops | 8.2 | 41.00 | 5.00 | 12.4 | 56.00 | 4.52 | ▼ −0.48 | | Vendor-B | reasoning-pro | Engineering | 77.5 | 5,811.00 | 74.98 | 191.6 | 12,527.00 | 65.38 | ▼ −9.60 | | Vendor-B | reasoning-mini | Engineering | 110.0 | 330.00 | 3.00 | 180.5 | 487.00 | 2.70 | ▼ −0.30 | | Vendor-C | open-weight-70b | Data Science | 220.0 | 264.00 | 1.20 | 410.0 | 451.00 | 1.10 | ▼ −0.10 | | Vendor-A | frontier-large | Customer Support | 15.2 | 456.00 | 30.00 | 19.1 | 497.00 | 26.02 | ▼ −3.98 | | Vendor-D | legal-domain | Legal | 2.1 | 420.00 | 200.00 | 2.4 | 504.00 | 210.00 | 🚨 +10.00 | > Note on data sparseness: frontier-mini (Revenue Ops), reasoning-mini (Engineering), open-weight-70b (Data Science), frontier-large (Customer Support), and legal-domain (Legal) have data only in April and

…output truncated at 250 words.

Step 2 output

# Forward-Looking Cost Model: Q4 2026 & Q1 2027 --- ## Anchor Points: September 2026 Run Rates All volume and price assumptions build from these actuals. | Vendor | Model | Team | Sep Tokens (M) | Sep Spend ($) | Implied $/M | Flag | |---|---|---|---:|---:|---:|---| | Vendor-A | frontier-large | Revenue Ops | 214.6 | 5,151 | 24.00 | 🚨 Agentic risk | | Vendor-A | frontier-mini | Revenue Ops | 12.4 | 56 | 4.52 | — | | Vendor-B | reasoning-pro | Engineering | 81.5 | 5,134 | 63.00 | — | | Vendor-B | reasoning-mini | Engineering | 180.5 | 487 | 2.70 | — | | Vendor-C | open-weight-70b | Data Science | 410.0 | 451 | 1.10 | — | | Vendor-A | frontier-large | Customer Support | 19.1 | 497 | 26.02 | — | | Vendor-D | legal-domain | Legal | 2.4 | 504 | 210.00 | 🚨 Price increase | > Monthly → Quarterly conversion: September actuals × 3 = Q4 base run rate before applying volume and price assumptions. This is the cleanest anchor given the within-Q3 acceleration already observed. --- ## Step 1 — Establish September Quarterly Run Rate (×3) | Vendor | Model | Team | Monthly Tokens (M) | Q Run-Rate Tokens (M) | Monthly Spend ($) | Q Run-Rate Spend ($) | |---|---|---|---:|---:|---:|---:| | Vendor-A | frontier-large | Revenue Ops | 214.6 | 643.8 | 5,151 | 15,453 | | Vendor-A | frontier-mini | Revenue

…output truncated at 250 words.

Step 3 output

# Renegotiation Brief — Vendor-A Account Review To: Account Manager, Vendor-A From: [FP&A / Procurement Lead] Date: October 2026 Re: frontier-large Contract Renewal — Pricing & Commitment Discussion --- Across frontier-large and frontier-mini, our year-to-date spend with Vendor-A totals approximately $35,100, with a forward run rate of $67,000–$84,000 over the next 12 months depending on workflow trajectory — making Vendor-A our single largest AI infrastructure vendor by a material margin. We want to grow this relationship deliberately and are prepared to make a formal volume commitment. Before doing so, we need to resolve three observations from our internal billing analysis: - Unit price is not reflecting our volume growth. Revenue Ops consumption on frontier-large grew from 42M tokens in April to 215M tokens in September — a 5× increase in six months — yet our blended rate moved only from $30.00/M to $24.00/M. Industry benchmarks and our own vendor comparisons suggest a account of this trajectory warrants a rate closer to $18–20/M at current scale. - Our Q4 volume commitment is substantial. Under our base forecast, Revenue Ops alone will consume approximately 966M frontier-large tokens in Q4 2026, with a full-year 2027 projection exceeding 3.5B tokens across all teams if current workflows scale as modeled. - The current trajectory creates budget pressure that will force architectural decisions. Our forward model shows a $16,000 quarterly swing between leaving workloads on frontier-large versus rerouting them to frontier-mini or an open-weight alternative at $1.10–$4.52/M. **We are asking for a blended frontier-large rate of

…output truncated at 250 words.

---

This content is for informational purposes only and is not financial, investment, or accounting advice. Verify outputs against authoritative sources before use.

Source: news.google.com

More for Finance professionals →

Get the next one in your inbox

One daily brief. Every story gets a hype verdict.

No spam. Unsubscribe anytime.

Exact prompts included · Untested steps are marked · Corrections are public