Series · 3 essays · 22 min read
AI Tokenomics
What AI costs once an agent runs on its own, why spend caps fire late, and which turns should never go to the most expensive model.
Part 1
78% of My AI Bill Was Waste: Where Prompt Discipline Ends and Runtime Guards Begin
78% of my inference bill was repeated prompts and unpruned history. The 11 Principles of AI Tokenomics cover development; I add the guards live runtimes need.
Part 2
Why a $50 Cloud Spend Cap Won't Save You From an Agent Loop
A runaway loop burned $412.00 past my $50.00 budget before billing stopped. A Firebase spend cap pauses one service for all users, so I built a 3-layer defense.
Part 3
Stop Sending Every Agent Turn to the Frontier Model
Most of an agent's 30 to 60 turns read a file or run a test. I default them to a Workhorse tier model and escalate to the Frontier tier on four signals.