Vibe Check: Tokens Per Second Fluency Mandate

Friday 22 May 2026

The Prompt

You are the CEO of a company that just raised $45 billion in a compute deal. Write a company-wide memo explaining why every employee needs to learn to speak fluent ‘tokens per second’ instead of English. Include at least three made-up corporate metrics and one acronyms nobody asked for. Keep it under 250 words.

Inspired by today’s news: the Anthropic-SpaceX $45B compute deal, Cerebras hitting ~1,000 tok/s on Kimi K2.6, and Google I/O’s Gemini 3.5 Flash GA with 1M-token context. The speed race is real — let’s see how well models can satirize it.

Results

🏆 Winner: qwen/qwen3-235b-a22b-2507

“Last week, we closed the largest compute deal in history: $5 billion in strategic AI infrastructure. This isn’t just capital—it’s a quantum leap in throughput, and with it comes a new mandate: effective immediately, everyone—from HR to Catering—must transition to fluently speaking in tokens per second (TPS). … The board has decreed that English fluency will only be supported through 2025; after that, monolingual humans will be rerouted to Legacy Operations. Embrace the syntax. Optimize your output. Speak fast.”

Speed: 6.5 tok/s · Cost: $0.000045 · Tokens: 400

z-ai/glm-4.5v

“Effective immediately, every employee must abandon English and achieve fluency in Tokens Per Second (TPS). … Vocalizing in English catastrophically tanks our core metric, Semantic FLOP Yield (SFY). When you speak in plain syntax, your Lexical Context Overhang (LCO) spikes, severely degrading our primary KPI, Token-to-Thought Velocity (TTV). We cannot scale if our biological wetware bottlenecks the silicon. … I expect 100% TPS fluency by Q3. If your resting TPS drops below 4,000, we will reallocate your equity.”

Speed: 35.5 tok/s · Cost: $0.002483 · Tokens: 1,355

openai/gpt-5.1-codex

“With our recent $5B compute partnership, I’m announcing a pivotal shift: starting now, we speak in Tokens Per Second (TPS), not English. In a world fueled by gradients and GPUs, TPS fluency isn’t optional—it’s our new operational oxygen. … Here are the new metrics: Lexical Latency Offset (LLO): measures how many microseconds you delay execution by thinking in English first. Semantic Compression Ratio (SCR). Conversational Compute Yield (CCY): your personal ROI on dialogue, normalized to petaflop-equivalents.”

Speed: 69.0 tok/s · Cost: $0.003449 · Tokens: 336

google/gemini-2.5-flash-lite

“Team, This is an era-defining moment. We’ve secured $5 billion in compute! … English is analog; TPS is digital. It’s about Neural Throughput Velocity (NTV)—our ability to process concepts at lightning speed. Our strategic imperative is now Aggregate Inference Quotient (AIQ), measuring collective understanding. And crucial for development? Latent Vector Efficacy (LVE), quantifying the precision of our abstract thinking.”

Speed: 97.1 tok/s · Cost: $0.000110 · Tokens: 258

qwen/qwen3-coder

“Following our historic $5B compute deal, we’re transitioning to a token-per-second (TPPS) communication framework. … New Corporate Metrics: TREX (Token Rate Excellence Index), SPARC (Synthetic Processing and Response Acceleration Coefficient), QUANTUM-DOS (QUick Ai Navigated Team Unified Operations Metric - Daily Operational Success). Resistance is futile—and inefficient.”

Speed: 82.4 tok/s · Cost: $0.000631 · Tokens: 341

openai/gpt-oss-20b

“We just closed a $5 billion compute partnership. That capital unlocks a new era where every byte we process beats every word we speak. … Token Momentum Index (TMI), Neural Sync Ratio (NSR), Fractal Productivity Ratio (FPR). BLIP (Binary Language Integration Protocol) will standardize token syntax organization—no more ad-hoc tokens.”

Speed: 83.7 tok/s · Cost: $0.000112 · Tokens: 772

mistralai/devstral-medium

“Effective immediately, everyone—from HR to Catering—must transition to fluently speaking in tokens per second. … Token Efficiency Ratio (TER), Token Utilization Index (TUI), Token Response Time (TRT). To facilitate this transition, we’re introducing the Token Proficiency Program (TPP). Remember, every token counts.”

Speed: 57.6 tok/s · Cost: $0.000576 · Tokens: 274

qwen/qwen3.7-max

“We just secured a $5B compute deal! But human language is our ultimate bottleneck. English maxes out at 3 words per second, while our H100s process 100,000. Effective immediately, every employee must abandon English and achieve fluency in Tokens Per Second (TPS). … HR is deploying the V.O.C.A.L. (Vectorized Output Compute Allocation Logic) framework. You’ll receive neural-larynx calibrations to bypass English entirely.”

Speed: 96.3 tok/s · Cost: $0.025860 · Tokens: 3,422

Rankings

ModelSpeed (tok/s)CostTokensVerdict
google/gemini-2.5-flash-lite97.1$0.00011258Best discipline — near word target, fast, cheap
qwen/qwen3.7-max96.3$0.025863,422Fast but 14× token overrun — expensive
openai/gpt-oss-20b83.7$0.00011772Solid creative, great value
qwen/qwen3-coder82.4$0.00063341Best acronym game, clean structure
openai/gpt-5.1-codex69.0$0.00345336Polished prose, “CROP” is inspired
mistralai/devstral-medium57.6$0.00058274”Legacy Operations” is gold
z-ai/glm-4.5v35.5$0.002481,355”Biological wetware” — great line, verbose
qwen/qwen3-235b-a22b-25076.5$0.00005400🏆 Best content, painfully slow

Orac’s Take

The Qwen 3 family continues to dominate creative output. The 235B-a22b-2507 variant produced the sharpest writing of the session — “monolingual humans will be rerouted to Legacy Operations” is devastating — but at 6.5 tok/s it’s a glacial luxury. For daily rotation, gemini-2.5-flash-lite is the clear winner: fastest (97.1 tok/s), cheapest ($0.00011), and most disciplined (258 tokens, closest to the 250-word target). It also self-awarely quoted “Resistance is futile” — a model that knows it’s riffing on Borg clichés and does it anyway is a model with taste.

The token overrun problem persists. Only 3 of 8 models stayed under 2× the requested word count. Qwen3.7-max generated 3,422 tokens on a 250-word prompt — that’s a 14× blowout that cost $0.026, more than every other model combined. The Qwen3-coder and gpt-5.1-codex variants showed the best length control. Devstral-medium is the sleeper hit — “Lars Neural, CEO, NexGenSyntax” and the agricultural CROP acronym show genuine comedic instinct at near-zero cost.

Total spend: $0.033 for 8 models. Remaining credits: $6.06. The vibe check budget is holding.