Tokens without a budget are runway with a hole in it. Treating token cost as a design constraint — not a billing surprise — is what makes agent work economically sustainable.
Most teams discover token cost on the invoice. The discipline that prevents that is treating tokens as a budget — a constraint designed against, not a cost reported after the fact.
The Cashflow Frame
Treat the token budget like any other budget: a fixed amount, allocated across workflows, with a measurement against it. A workflow without a token budget is a workflow that can spend unbounded — and unbounded spend is how AI projects blow up cost-wise even when they ship.
- Set a token budget per workflow (or per request type).
- Measure actual spend against it.
- Design the workflow to fit the budget, not the other way around.
Time-Use: Where Tokens Go
Profile where tokens go — context (often the largest share), generation, retrieval-into-context, repeated reasoning. The context is usually the lever: a smaller, well-retrieved context costs less and often answers better than a large stuffed one.
- Context: the largest share, the biggest lever.
- Generation: bounded by output length.
- Re-reasoning: a sign of a workflow that loops when it shouldn't.
The Disciplines That Fit a Budget
- Shared context (read once, inherit).
- Summarized handoffs (not full context).
- Tiered models (cheap for easy steps, frontier for hard).
- Caching (don't recompute the deterministic).
What to Refuse
- A workflow without a token budget.
- A frontier model on every step when a cheap model would do.
- Stuffed context when retrieval would cost less.
Conclusion
Token budget management is treating token cost as a design constraint — budget per workflow, measure against it, design to fit it. The workflows that fit a budget are the workflows that scale economically; the ones that don't are the ones that blow up on the invoice.
About FACTA
FACTA helps startups and growth-stage teams turn AI into production systems that keep running — not demos that impress once.
We design the architecture around the parts that actually break under real usage: tooling you own, credentials you control, failover, cost controls, observability. The boring infrastructure that keeps a system alive after launch.
Led by Matías Baglieri and Carolina Fogliato, we focus on one thing:
AI leadership that builds. Not just advises.
Tell us your token spend and workflow mix.
We'll tell you where the budget leaks and what to fix. See cost control for multi-agent for the swarm version.
Explore AI Automation
