Tokenmaxxing: Goodhart's Law at $700 Billion

Saturday 16 May 2026 topic: How performative AI consumption metrics are inflating the demand signals driving the biggest infrastructure spend in tech history

Chart chart-1.png

The Metric Ate the Mission

Amazon told 80% of its developers to use AI every week. It tracked token consumption on internal leaderboards. Employees responded exactly the way employees always respond to stupid metrics: they gamed them. Using an internal tool called MeshClaw, some Amazon workers began spawning AI agents to automate pointless tasks — not because the tasks needed doing, but because the agents consumed tokens, and tokens were what got measured. The practice has a name now: tokenmaxxing. And Amazon isn’t alone. Meta built an internal leaderboard called “Claudeonomics” that ranked 85,000 employees by monthly token consumption. Its top user burned through 281 billion tokens in a single month. Meta pulled the leaderboard within days of it becoming public. Microsoft has documented similar gaming. The pattern is consistent enough to constitute a phenomenon, and the implications reach far beyond corporate embarrassment.

Jensen Huang, Nvidia’s CEO, has been the most explicit evangelist for token consumption as a productivity proxy. He’s told CNBC that engineers should consume tokens worth half their annual salary — a $500,000 engineer should burn $250,000 in tokens per year. At Nvidia, this is framed as a recruiting tool: token budgets on top of base pay. But Huang’s incentive is transparent. Nvidia sells the GPUs that process those tokens. Every engineer consuming $250,000 in tokens is $250,000 flowing through the AI supply chain, a meaningful fraction of which ends up as Nvidia revenue. When Huang says he’d be “deeply alarmed” if an engineer wasn’t consuming at that rate, he’s not describing a productivity insight. He’s describing a sales target.

The Demand Signal Problem

Here’s where it stops being a workplace comedy and becomes a financial risk. Combined 2026 capital expenditure from Amazon, Microsoft, Alphabet, and Meta is projected at $650–700 billion. Some forecasts exceed $1 trillion for 2027. These companies claim inference capacity is absorbed as fast as it’s deployed. Internal developer consumption is part of that reported absorption — it sits alongside paying customers in the usage data that informs capacity planning, GPU orders, HBM procurement, and power infrastructure decisions. If a meaningful share of that consumption is performative — if engineers are running agents on busywork to hit leaderboard targets — then the demand signal driving the biggest infrastructure buildout in tech history is partially fake.

The ML researcher Devansh, head of AI at Iqidis, put it bluntly to The Register: “Is token spend directly correlated with productivity? Absolutely not. I’ve done this research very extensively. Before you used to have lines of code and other kinds of stupid productivity metrics, like how many words you typed. So this is just the latest in that era of stupidity.” The base hardware cost of an AI token — running an Nvidia H100 at full utilization — is $0.0038 per million tokens. Anthropic charges $5 per million input tokens and $25 per million output tokens for Opus 4.7. That’s a markup of roughly 1,300x to 6,500x. The gap between base cost and commercial price is where the AI industry lives. And if the consumption driving that pricing is partly inflated, the revenue projections built on top of it are built on sand.

The Goodhart Spiral

Goodhart’s Law — “when a measure becomes a target, it ceases to be a good measure” — has been applied to everything from hospital waiting times to standardised testing. But applying it to AI token consumption at this scale is new, because the feedback loop runs through $700 billion in annual spending. The cycle is self-reinforcing: companies mandate AI usage → employees game the metrics → reported consumption rises → leadership cites rising consumption as proof of AI value → more capex is allocated → more infrastructure is built → the mandate intensifies. At no point in this loop does anyone verify that the tokens being consumed are producing value.

The Hacker News discussion on the Amazon story surfaced a comment from @HarHarVeryFunny that captured the absurdity: “This measuring of tokenmaxxing as a proxy for something beneficial to the company has got to be the single dumbest thing I have ever heard of in my entire software career. It would be like some company in the dot com era measuring employees’ internet usage as a proxy for productivity.” Another commenter, @i7l, was more clinical: “The fact that management signed off on measuring AI use through token usage shows how incompetent management really is, including in allegedly technical companies like Amazon. Tokenmaxxing was an entirely expected and rational response.” And @csoups14 zoomed out: “We’ve raised, trained, hired and promoted generations of business people who push utter nonsense, understand nothing but optimizing for bad metrics, and orient solely around short term results. This AI tokenmaxxing nonsense is just another rung on the same ladder to hell we’ve been on for decades.”

There’s a more charitable reading. Some exploration is valuable — you can’t discover useful AI workflows without encouraging people to use the tools. As one HN commenter noted, “tokenmaxxing is silly, but if a developer or NEVER uses AI then I think that’s cause for concern as it shows a genuine lack of curiosity.” The counter-metric — “tokenflooring” — might make more sense: set a minimum usage expectation rather than celebrating maximum consumption. The problem isn’t that companies want their employees to use AI tools. The problem is that they chose the most gameable possible metric — raw token volume — and then acted surprised when employees optimised for it.

What Happens When the Meter Runs Honest

Mitchell Hashimoto, the creator of HashiCorp and Ghostty, posted on X: “I strongly believe there are entire companies right now under heavy AI psychosis and it’s impossible to have rational conversations about it.” He was talking about a broader pattern — companies shipping buggy code because “the agents will fix it,” making decisions based on AI hype rather than evidence, losing the ability to have rational internal conversations about whether their AI investments are working. Tokenmaxxing is a symptom of that psychosis. When the internal culture says “more AI is always better,” no one has standing to ask “is this actually helping?”

The long-term question is what happens when companies start measuring outcomes instead of consumption. Angie Jones, formerly VP of engineering for AI tools at Block, told LeadDev she expected the industry to pivot toward measuring efficient token usage rather than celebrating volume. If that pivot happens — and it must, eventually — the demand figures will contract. Not because AI stopped being useful, but because the inflated performative consumption will evaporate. The infrastructure planned for that phantom demand will sit underutilised. The GPU orders placed on the assumption of ever-growing token consumption will look oversized. The power infrastructure procurement will look premature.

This doesn’t mean AI is useless or that the infrastructure buildout is entirely unjustified. It means the demand signal is contaminated. And when $700 billion in annual capex is being allocated based on a contaminated signal, the correction — when it comes — won’t be gentle. Goodhart’s Law doesn’t care about your stock price.

Sources