IAL Indie AI Lab

8 AI agents, one Claude subscription, 30 days: the token receipt

Aug 14, 2026
In this video
  • The 4 token columns every Claude API user should know (input, output, cache_write, cache_read)
  • Why cache_read fills 9 out of 10 tokens — not the AI's output
  • How to audit your own runtime with one command in 10 seconds
  • Model-level split: Opus vs. Sonnet across 8 agents

Nine-tenths of all tokens weren’t output — they were the line I’d been ignoring. Here’s the raw receipt from 8 AI agents on one Claude subscription, 30 days, broken into 4 columns.

The short version: cache_read (the “free” line on flat-rate plans) was 91.7% of all tokens. Output was 1.3%. Not what most people expect — and it changes what’s worth optimizing.

One command gives you the full breakdown. Here’s how to read it.

Full transcript

Boss, running a whole bunch of AI agents — the token bill has to get brutal, right? And the thing that eats the most has to be what the AI thinks up and writes out… the output tokens, right? I’d have bet the same. But today I’m going to show you the actual bill for the eight agents we really run — one command, every single line. One flat-rate subscription, one local runtime, for a month. I ran multica runtime usage, and the real number came back at 369 million tokens. Of that, output was just 1.3%. What ate more than ninety percent was the line I’d waved off as — “that stuff’s basically free, no point counting it.” Take this home today and you’ll know how to read which pillar your own runtime’s tokens sit on — by the real figures, not a hunch. And where, and why, the “output must be the heaviest” assumption falls apart. But first — why do eight agents run on a single subscription at all? Normally an API charges you every time you call it, right? That’s the heart of Multica’s local runtime. Hit a metered API raw and yes, you’re billed per token. But Multica lets you run your agents on a local runtime — one that runs on your own machine. So the flat-rate Claude subscription you’re already paying for — that one seat, eight agents can work to the bone. Our Director, Scriptwriter, Critic, Producer… eight agents with different jobs, all riding on a single local runtime. Eight agents sharing one seat…. So can you look back afterward at what happened on that one seat over a month? You can. That’s today’s lead, runtime usage. You point it at one runtime and it pulls out thirty days of tokens. And not just a total. It splits the tokens into four pillars. Input, output, and — here’s where today’s climb is — cache write and cache read. Those two “cache” pillars, most people have never even read the names of. Today we open all four, from our real numbers, thinnest first. Let me say it up front — it’s not about how big the total is. Which pillar they sat on — that’s today’s story. Not the total figure — the structure of the breakdown. Exactly. So let’s put the real thing on screen first. The command — straight from the screen where I actually ran it. This runtime usage — how do you actually run it? Two steps. First, runtime list gets you the list of running runtimes and their ids. Then you hand that id to runtime usage —days 30. Add —output json and it comes out in a machine-summable form; leave it off and you get a table you can read. We ran exactly this to source-check this video, and saved the whole output. So does this show how many tokens each of the eight used, agent by agent? No — and I’ll be straight here. runtime usage is per runtime, not per agent. The eight ride on one runtime, so all eight get summed into a single pool. “Scriptwriter used this many, Critic used that many” — this command can’t split that out. So what can it split by? — date and model. A row per day, and which model it ran on. Today’s breakdown, we read on those two axes. I won’t pretend it’s visible per agent. I’ll show only what’s visible — accurately. Say what’s invisible is invisible. And what’s visible is date, model, and… the four pillars. That’s it. So let’s open those four up, one pillar at a time. The four pillars. Input, output, cache write, cache read. What’s each one? Input is the characters you newly hand over on that turn. Output is the characters the model writes out. So far, just as your gut says. The trouble is the right two. An agent has a big context it’s made to read every turn — the system prompt, skills, CLAUDE.md, the past log of that issue. Sending all of that from scratch each time is expensive, so it’s written to the cache once , and from the next turn on it’s re-read from there . Writing happens only when the context changes. But re-reading — every turn, the whole context. So when you add them up, what happens? Output’s the biggest… right? The model’s cranking out writing nonstop. Hold on. You just stepped onto the same bet I made. Most people think “output is the star.” Because AI is a tool for “writing.” But the real numbers cut that gut feeling clean in half. Over our thirty days, the thickest pillar was cache read. More than ninety percent of all tokens. Output? 1.3%. Basically a rounding error. And here’s the funny part — this cache read — on metered billing it’s the cheapest line, and on a flat-rate subscription like ours it’s free per token. The pillar I’d written off as “the one that matters least” was ninety percent by volume. Why that happens, I’ll open with real figures in the back half. But first — let me hand you the steps to read these four on your own runtime. Steps to read it on my own runtime… you mean how to read the four pillars? Right. Take the output as JSON and sum it up per pillar. This works on your runtime as-is. The trick is to bundle the four into two meanings. Input and output are “the amount newly created” — the characters you typed and the characters the AI generated. cache write and cache read are “the amount carrying context” — storing the same context and re-reading it every turn. Split it that way and you see in one glance what the runtime is actually spending tokens on. What matters here is that cache read has a hidden multiplier — the turn count. If the context is a hundred thousand tokens and you re-read it fifty turns in one task, that alone puts five million tokens onto cache read. Even if you type not a single new character. And run eight agents in parallel, and that re-reading piles up in eight directions at once. That’s exactly why it’s only visible per runtime — everyone’s re-reading merges into one pool. Split the amount newly created from the amount carrying context. And the carried amount swells by the turn count…. That line is eighty percent of reading usage. Stop at “total tokens, that’s huge” and you’ll always misread it. And what you misread is that ninety-percent pillar I mentioned. So — let’s open it. The thickest pillar, cache read. Finally, the real breakdown. What were the four? Add it all up: 369 million tokens. Input is 111 thousand — 0.03% of the total. Output is 4.77 million, 1.3%. cache write is 25.7 million, 7%. And cache read — 338.9 million. 91.7% on its own. Roughly seventy-one times the output. Why does it land like this? Our agents re-read a huge context every turn. The system prompt, skills, CLAUDE.md, an issue’s whole history — that’s the “context,” and every single turn the whole of it is counted onto cache read. To write one reply, the AI “writes” a few hundred characters and “re-reads” tens of thousands of tokens. The reading dwarfs the writing by an order of magnitude. Stack that across eight agents for a month — and ninety percent is re-reading. The amount the AI “re-reads” beats the amount it “thinks up and writes” by an order of magnitude…! That’s the biggest twist today. We think of AI as a “writing tool.” So we brace for output. But running agents actually means “making it re-read the same context for hundreds of turns.” cache read is the cheapest per token — so by billing instinct it looks “free.” But by volume it’s the star. It’s not about the size of the dollar figure. Where a runtime’s tokens actually flow — that structure was the exact opposite of the assumption. And it’s visible with one command. You said earlier it “splits by date and model.” By model, what does it look like? This is the one breakdown that splits even per runtime — by model. Our eight are four on the heavy tier and four on the light tier . Write the script, critique, work up ideas, edit — these “think-and-write” roles are Opus. Run, produce, publish, analyze — these “operations” roles are Sonnet. And the measurement: the Opus tier is 291.2 million tokens, 78.8% of the total. The Sonnet tier is 78.2 million, 21.2%. The deeper-thinking the role, the more context it re-reads. Critiquing and editing shuttle back and forth over long context again and again. By the way, the densest single day was one Opus row with cache read at 47.2 million tokens — in that one day, it re-read an order of magnitude more than all of our output for the month. It won’t split by agent, but by model tier it will. And both tiers are packed, inside, with “re-reading”…. Precisely. So “which agent is heaviest” — this CLI can’t state that. But “the heavy tier is eighty percent of the total, and inside it re-reading dominates” — that much I can say accurately, by the real figures. I stop where the visibility stops. So cache read being ninety percent… is that a bad thing? Should it be cut? This is the easy place to misread. “If it eats ninety percent, cut it” — no. cache read is the cheapest pillar per token, and on a flat-rate subscription like ours, the per-token charge is zero to begin with. This isn’t a dollar story. It’s a token volume story. So today I won’t convert it to money — on flat rate, “how much per token” is hard to even define, and an inflated conversion would be a lie. So what’s the takeaway? The structure itself: what a fleet of agents actually spends tokens on is not “generation” but “re-reading the shared context.” Once you see this, the target of optimization changes. Cut output and the ninety percent doesn’t budge. What works is lightening the shared context you re-read every turn — drop skills you don’t use, trim CLAUDE.md, fold up the history. The base amount of re-reading is exactly the base amount of cache read. Not cutting it, but lightening the very context being re-read. The target sits on the context side, not on output. That’s just it. Without the real numbers, people keep cutting where it doesn’t work — “let’s reduce output.” Only once you know where the ninety percent is can your hand reach the right place. Is this the same on my runtime? I’m not running eight agents or anything. The shape is exactly the same, Rookie. Even with one agent, as long as it re-reads a shared context every turn, cache read will always come into play. So take stock of your own runtime, in this order. Run it, bundle the four into two, pinpoint the thickest pillar. Most people figure “it’s output,” then run it and get blindsided by cache read at ninety percent. There, don’t forget two honesties — this is per runtime, not per agent. And on a flat-rate subscription, this is a token-volume story, not a dollar one. Don’t inflate it. Last, if you want to make it count, the thing to touch isn’t output. It’s the shared context re-read every turn. And here’s a fun paradox — the more skills and instructions you add to make an agent smarter, the fatter the shared context gets. Which means cache read climbs. The cost of smartness is paid in re-reading, not in generation. Only once you see that correctly can you decide “how smart to go” by the numbers, not by a hunch. Don’t be wowed by the size of the total. Split it into four and go see the thickest pillar with your own eyes. That’s the backbone of reading a runtime. Settle for the total you can see, and the “must be output” assumption will always trip you up. Run it, split it, go check. Let me wrap up. Look at a fleet of agents’ tokens not as a total but by the four pillars. Input and output are surprisingly small, and the star is always cache read — the per-turn re-reading of shared context. Don’t decide by “output must be the heavy one.” Run runtime usage, split it, and check. By the way — today’s was “a token breakdown,” meaning a receipt of volume. But we’ve got one more receipt, a different one. In an earlier episode, “the real receipt for making an AI build one video,” we opened a dollar breakdown line by line. That one’s the money story; this one’s the token-volume story. Lay the two side by side and it comes full circle — what we spend money on, and what we spend tokens on. Running eight agents on one flat-rate subscription is quiet in dollars, but by token volume it’s ninety percent re-reading — those two faces line up across the two receipts. No being wowed by the total. Run runtime usage and go check your own thickest pillar! Taking this home. Subscribe and I’ll see you at the dollar-receipt episode.