Series · 3 essays · 22 min read

AI Tokenomics

What AI costs once an agent runs on its own, why spend caps fire late, and which turns should never go to the most expensive model.

Start with Part 1

  1. Part 1

    78% of My AI Bill Was Waste: Where Prompt Discipline Ends and Runtime Guards Begin

    5 min read

    78% of my inference bill was repeated prompts and unpruned history. The 11 Principles of AI Tokenomics cover development; I add the guards live runtimes need.

  2. Part 2

    Why a $50 Cloud Spend Cap Won't Save You From an Agent Loop

    7 min read

    A runaway loop burned $412.00 past my $50.00 budget before billing stopped. A Firebase spend cap pauses one service for all users, so I built a 3-layer defense.

  3. Part 3

    Stop Sending Every Agent Turn to the Frontier Model

    10 min read

    Most of an agent's 30 to 60 turns read a file or run a test. I default them to a Workhorse tier model and escalate to the Frontier tier on four signals.