When the LLM invoice spikes, teams open the cloud console. The bigger lever is usually the product: which actions call a model, how much context you stuff into prompts, whether the UI invites retries, and whether tenants have caps. Infra tuning helps. Product design decides the shape of demand. Founders who treat spend as “an infra line” keep shipping UX that burns tokens for low-value moments — then wonder why margin and pricing never meet.
UX patterns that quietly burn money
Autocomplete that fires on every keystroke. “Regenerate” without rate limits. Chat that resends the entire history plus a vault of retrieved chunks on each turn. Agents that narrate every thought with a frontier model. None of these look like waste in a design review; all of them show up as COGS. Design for deliberate AI moments. Prefer explicit user intent for expensive calls. Debounce. Cache identical prompts. Stream smaller drafts before offering a premium rewrite. Make regenerate cost-aware in the UI when appropriate — especially on lower-priced plans.
Route models by task class
Not every step needs the smartest model. Classification, extraction, formatting, and simple rewrites often work on smaller or cheaper models. Reserve frontier models for hard reasoning and customer-visible writing where quality differences matter. Product owns the routing policy with evals. Engineering implements it. If routing is ad hoc, spend drifts and quality becomes inconsistent across the same feature.
Cache and compress context
Prompt caching, embedding caches, and normalized retrieval beats drowning the model in duplicated boilerplate. Summarize long threads into structured state instead of replaying megabytes of chat. Smaller, better context often beats larger, noisier context on both cost and quality.
Per-tenant caps are a feature
Caps protect you and teach customers how the product is meant to be used. Soft warnings, plan-aware limits, and clear upgrade paths beat surprise invoices. Show usage in-product so champions can manage their own teams — otherwise they discover spend when finance forwards a bill.
Instrument product-shaped cost
Attribute tokens to feature, tenant, and user journey step. Pair cost with outcome metrics you already track — completed workflows, retained users — without inventing ROI fairy tales. When product managers see cost beside engagement, they stop treating models as free magic. The LLM bill will always have an infra component. The part you can redesign weekly lives in product and design. Start there.
Plan packaging around product cost centers
If a feature is structurally expensive, it belongs on a higher plan or behind credits — not hidden inside an all-you-can-eat tier that finance will regret. Product managers should estimate token cost per happy-path use during design review, the same way they estimate engineering weeks. Kill or redesign features whose cost per retained user cannot work even after routing and caching. Sentiment for a flashy demo is not a business case.
Team workflows that control spend
Weekly cost review with product, eng, and finance beats quarterly shock. Celebrate cost-per-success improvements like feature launches. Add budget alerts to Slack for anomalies. Make “why did this feature get expensive” a blameless inquiry into UX and routing — not a hunt for a villain. Give design partners visibility into their usage early. Customers who can self-manage become collaborators on efficiency instead of adversaries on invoices.
Quality and cost are joint OKRs
Optimizing only for cost produces useless cheap answers. Optimizing only for quality produces beautiful insolvency. Joint OKRs — task success within a cost envelope — keep the org honest. Evals should include a cost dimension so prompt changes that double spend for a tiny quality bump are visible before merge.
Design reviews that include a cost sketch
Add a lightweight cost sketch to feature specs: expected calls per user session, context size class, model tier, and rough monthly cost at target usage. It will be wrong at first — that is fine. The point is to force the conversation before engineering paints the feature into a corner. Reject specs that say “just call the best model” with no fallback. Require a degradation path: smaller model, extractive answer, or human handoff when budgets trip. Designers should see cost as a constraint like accessibility and performance. Constraints create better craft. Unlimited magic produces less thoughtful interfaces and worse bills.
Procurement and vendor management also belong in the product cost story. Negotiate rate cards and commitments with model providers based on measured mix of task classes — not on hope. Share forecasts grounded in product roadmaps so finance is not blindsided by a new agent feature that multiplies calls overnight. When leadership asks “why is AI so expensive,” answer with feature-level attribution and a redesign plan. That posture turns a scary invoice into a managed product portfolio decision.
Leave the team with a simple rule: no new AI surface ships without a cost sketch, a cap strategy, and a routing default. Infra will still negotiate rates and tune clusters — but product owns whether the bill makes sense. That ownership is how AI features stay in the portfolio instead of becoming the reason margin reviews turn ugly.



