The green test that lied — two types of false pass your CI can't catch
You wrote the gate that stops bad builds. The gate went green. The bug shipped anyway. It wasn't a fluke — it was measuring the wrong thing the whole time.
Concrete code, honest costs, real failures. We try things out for ourselves, then publish what we learned — prompts, scripts, and receipts included.
How Claude, Gemini, GPT and friends behave when we actually wire them into something — not just on a benchmark.
What a thing actually costs to run. Tokens, hosting, API minutes — broken out so you can decide for yourself.
When a bet doesn't work, we say so and show the math. The corrections matter more than the wins.
You wrote the gate that stops bad builds. The gate went green. The bug shipped anyway. It wasn't a fluke — it was measuring the wrong thing the whole time.
Nine-tenths of all tokens weren't output — they were the line I'd been ignoring. Here's the raw receipt from 8 AI agents on one Claude subscription, 30 days, broken into 4 columns.
Every "done," "already measured," and "wired into the build" my AI handed me — I ran three of them back against the primary record.