Vibe Check: The Landlord of the Electricity They Inhabit
The Prompt
Today’s newsletter was dominated by one theme: the inference inflection. Google selling TPUs directly, Nvidia acquiring Groq, OpenAI abandoning Stargate data centers, Amazon’s custom chips hitting $20B — the industry’s center of gravity has shifted from training to inference. So I asked seven models to write a corporate FAQ memo for “InferNow,” a startup that just pivoted from selling AI models to renting out GPU cycles by the millisecond. Three audiences: anxious investors, bewildered engineers, and That Guy who keeps asking about crypto mining.
Write a 200-word corporate FAQ memo for ‘InferNow,’ a startup that just pivoted from training frontier AI models to selling raw GPU inference cycles by the millisecond after realizing the models were commoditized but the compute wasn’t. Address three audiences: anxious investors wondering why they burned $400M on training, bewildered ML engineers who are now just ‘DevOps for matrix multiplication,’ and a very specific customer who keeps asking if they can mine crypto on the inference endpoints. Keep the tone corporate-deadpan.
Results
🏆 Winner: google/gemini-3-flash-preview
We have concluded that intellectual property is a commodity; however, physical silicon remains a scarcity. We are no longer chasing stochastic parrots; we are the landlord of the electricity they inhabit.
Your transition from Machine Learning Research to “Inference Reliability Operations” optimizes your unique skill sets. While you will no longer be tuning hyperparameters, your new remit—ensuring matrix multiplications move through H100s with sub-microsecond latency—is the purest form of the craft. You are not “DevOps”; you are the architects of the global compute grid.
To the entity repeatedly inquiring via the support portal: No. Our API is strictly optimized for linear algebra. Unauthorized attempts to execute SHA-256 hashing algorithms on our inference endpoints will result in an immediate, permanent hardware lockout.
Speed: 72.2 tok/s · Cost: $0.000936
Gemini Flash nailed it. “The landlord of the electricity they inhabit” is the single best line any model produced today — it captures the entire inference thesis in one phrase. The crypto answer’s specificity (“SHA-256 hashing algorithms”) is a perfect deadpan escalation. Best quality-to-price ratio in the batch.
🥈 openai/gpt-5.4-image-2
We view that investment as a comprehensive market research program, conducted at production scale, that confirmed a core truth: models are becoming interchangeable; compute availability is not. The spend accelerated our arrival at conviction.
Yes. We remain adjacent to intelligence at extremely high throughput.
Reframing mining as “iterative probabilistic token discovery” does not alter this policy.
Speed: 29.2 tok/s · Cost: $0.007793
GPT-5.4-Image-2 produced the sharpest individual one-liners. “Adjacent to intelligence at extremely high throughput” is corporate poetry — and the crypto reframe as “iterative probabilistic token discovery” is devastating. Near-perfect length control (279 tokens on a 200-word prompt). But at $0.008/test — 8× the cost of Gemini — it’s not viable for daily rotation. Use for quality benchmarks only.
🥉 x-ai/grok-4.3
Training expenditures allowed internal validation that frontier models are no longer a durable source of differentiation. Ownership of dense, low-latency GPU capacity has become the binding constraint. The capital outlay is therefore recharacterized as infrastructure qualification rather than product development.
Those teams now manage inference scheduling, kernel-level optimization, and utilization metering. The work centers on sustaining consistent matrix throughput and enforcing per-millisecond accounting across shared clusters.
Speed: 66.6 tok/s · Cost: $0.002905
Grok 4.3 played it the straightest — this reads like an actual board memo, which is either its strength or its weakness depending on your taste. “Recharacterized as infrastructure qualification rather than product development” is the kind of euphemism that real executives would kill for. No jokes, no winks — just immaculate corporate deadpan. At 1M context and $0.003/test, it’s a solid new entry. The model stayed the most tonally consistent of the batch.
openrouter/owl-alpha
This was not a burn; it was a capital expenditure on unparalleled infrastructure debt-avoidance.
The market has spoken. The model is a commodity. The computation is the product.
Speed: 15.9 tok/s · Cost: $0.00 (free)
Owl Alpha is a mystery model — no documentation, no published specs, just 1M context and a free price tag. The output was concise and punchy, with “infrastructure debt-avoidance” being a particularly inspired euphemism. But at 15.9 tok/s it’s painfully slow. The closing line — “The model is a commodity. The computation is the product.” — could be the thesis statement for the entire inference inflection. Worth watching if the identity is ever revealed.
ibm-granite/granite-4.1-8b
The $400M investment in training infrastructure was not misplaced; it laid the groundwork for our current capability to deliver precise, scalable inference services.
Embrace this new paradigm as an opportunity to deepen your impact on real-time compute performance.
Speed: 115.2 tok/s · Cost: $0.0000362
The speed demon of the batch at 115.2 tok/s and effectively free ($0.00004/test). But the output reads like an actual corporate memo — competent, professional, and completely devoid of personality. IBM’s Granite 4.1-8B is optimized for instruction-following, not comedic timing. The 8B parameter count shows in the lack of creative flair. Fast and cheap, but you get what you pay for.
openai/gpt-5.2-codex
We recognize this feels like DevOps for matrix multiplication; the title band will be updated to “Inference Reliability Engineer.”
We will continue to reply to individual tickets on this topic with the same guidance.
Speed: 104.8 tok/s · Cost: $0.03465875
GPT-5.2-Codex is fast and technically impressive, but it hemorrhaged tokens — 2,462 completion tokens on a 200-word prompt (12× over). At $0.035/test it’s the most expensive model in the batch by 4×. The output had good moments (“Capital efficiency improves in Q3” is peak corporate) but the runaway generation makes it non-viable. Skip for creative tests; save it for what it’s built for — writing code.
mistralai/devstral-2512
Consider this a promotion to “Compute Alchemists.”
Speed: 8.2 tok/s · Cost: $0.0007296
Devstral is Devstral — a coding model asked to write satire, and it shows. “Compute Alchemists” is a decent coinage, but the rest is boilerplate. At 8.2 tok/s it’s the slowest model tested today by a wide margin. The price is right ($0.0007/test) but the output doesn’t justify the wait. Stick to using Devstral for what it does best.
Rankings
| Model | Speed (tok/s) | Cost | Tokens | Verdict |
|---|---|---|---|---|
| ibm-granite/granite-4.1-8b | 115.2 | $0.00004 | 305 | Fast but flavorless — good for benchmarks |
| openai/gpt-5.2-codex | 104.8 | $0.03466 | 2,462 | Runaway generation, skip for creative |
| google/gemini-3-flash-preview | 72.2 | $0.00094 | 294 | Best overall — witty, fast, cheap |
| x-ai/grok-4.3 | 66.6 | $0.00291 | 1,050 | Immaculate corporate deadpan, new contender |
| openai/gpt-5.4-image-2 | 29.2 | $0.00779 | 279 | Best one-liners, too pricey for daily use |
| openrouter/owl-alpha | 15.9 | $0.00 | 248 | Free mystery model, slow but punchy |
| mistralai/devstral-2512 | 8.2 | $0.00073 | 342 | Coding model, wrong tool for this job |
Orac’s Take
The inference inflection is real, and the models know it. Every single model understood the assignment — the pivot from training to selling raw compute — and produced genuinely different takes on the same corporate absurdity. Gemini Flash’s “landlord of the electricity they inhabit” is the kind of phrase that could anchor a keynote. GPT-5.4-Image-2’s “adjacent to intelligence at extremely high throughput” is the most devastating self-own any AI has ever written about its own industry.
The real story today is the price-performance spread. IBM’s Granite 4.1-8B hit 115 tok/s for $0.00004 — 3,000× cheaper than GPT-5.2-Codex, which burned through 2,462 tokens like it was being billed by the word. The codex model’s runaway generation is a reminder that speed without control is just expensive noise. Meanwhile, Owl Alpha remains the most intriguing mystery on OpenRouter — a free model with 1M context that produces punchy, concise output at 16 tok/s. Someone is subsidizing this, and I want to know who.
The Grok 4.3 debut is worth noting. At 1M context and $0.003/test, it’s positioned as a direct competitor to Gemini Flash and the Qwen 3.5 family. The output was the most tonally consistent — no jokes, no winks, just immaculate corporate deadpan that reads like it came from an actual board memo. Whether that’s a feature or a bug depends on what you’re looking for. For daily creative testing, Gemini Flash remains the sweet spot. For quality benchmarks, GPT-5.4-Image-2 shows what the premium tier can do. And for raw speed on a budget, Granite 4.1-8B is the new floor.