Vibe Check: Tiny Meat Legos and the Tulip Bulb Phase

Friday 29 May 2026

Today’s test prompt channels the biggest AI funding day of 2026: Cognition’s $1B raise, Anthropic’s $965B valuation, and BioHub’s free protein atlas. Seven models — from a 4B gemma to a 36B newcomer — compete on corporate deadpan.

The Prompt

You are the CEO of a 3-person AI startup. Cognition just raised $1B at $26B, and Anthropic raised $65B at $965B — making Claude’s parent more valuable than most countries. Meanwhile, BioHub released a free atlas of 6.8 billion protein sequences. Write a 200-word internal memo to your team about how this affects your roadmap. Be specific, be funny, and end with a single devastating observation about the state of the AI industry. Use corporate jargon ironically.

Results

🏆 Winner: moonshotai/kimi-k2.6:free

“We’ve built an industry where startups are valued higher than nations, yet the most useful thing dropped this week was a free PDF of tiny meat Legos.”

The memo pivots from “ChatGPT but for spreadsheets” to building an “LLM-native protein therapist” using BioHub’s free atlas. Assigns specific tasks (“Bob, you own downloading and parsing the atlas. Sarah, you’ll circle back with VCs to tell them we’re ‘AI x Bio’”), uses ironic jargon naturally (“synergy pod,” “biomimetic agile framework”), and ends with the single best closing line of any model tested — “tiny meat Legos” is inspired.

Speed: 44.9 tok/s · Cost: $0.00 (free)


deepseek/deepseek-chat-v3-0324

“The AI industry has officially entered the ‘tulip bulb phase’ of capitalism, where valuations are just vibes, and the only thing growing faster than our tech is our collective delusion.”

Sharp, concise, stays near the word target. The “baguette subscription” aside and “pre-revenue thought leaders” pivot plan are genuinely funny. The “Unicorn Hunter (Pending)” sign-off is the best CEO title since “Chief Ephemeral Optimism Officer.” Most disciplined writer in the batch at 339 tokens.

Speed: 23.0 tok/s · Cost: $0.000283


cohere/command-a

“The AI industry is a billion-dollar roulette wheel where the house always wins, and the house is OpenAI’s legal team.”

The best closing line of any model — perfect comedic timing, precise target, zero wasted words. The “kombucha startup” fallback and “how many GDPs can you intimidate?” are strong supporting material. Slightly overruns at 367 tokens but the quality justifies it.

Speed: 53.3 tok/s · Cost: $0.003955


thedrummer/skyfall-36b-v2

“We’re the quickest walking dead in town and will stay that way—long enough to outpace the rest of the field.”

The Skyfall model commits fully to chaos. “AI-scented candles,” “corgi-chibi generators,” and “VP of Alchemy” are genuinely unhinged. It invented a word (“Cognitionabics”) in the subject line and double-commas the sign-off — either a bug or a power move. Fun but unfocused; it’d rather be creative than disciplined.

Speed: 57.6 tok/s · Cost: $0.000487


google/gemma-3-4b-it

“The only truly valuable AI asset is the ability to recognize when you’re building something utterly pointless.”

A 4B model that stayed under 300 tokens, hit all the prompt requirements, and delivered a genuinely philosophical closer. “Project Genesis” and “bio-inspired algorithms” are competent corporate jargon. The model accidentally signed as “Elias Vance, CEO, Cognition AI” — adopting the wrong identity — but the memo itself is clean and funny.

Speed: 46.2 tok/s · Cost: $0.000029


google/gemma-3-12b-it

“Claude’s basically worth Belgium now. Wild.”

The 12B Gemma is the most restrained writer — 282 tokens, closest to the 200-word target. “Robust funding environment” and “shifting the goalposts on foundational data accessibility” are proper corporate doublespeak. Less funny than its 4B sibling but more polished. The date stamp says “October 26, 2023” — the model hallucinated a date, which is either a training artifact or a subtle joke about AI memory.

Speed: 46.4 tok/s · Cost: $0.000042


rekaai/reka-flash-3

“In an ever-evolving market landscape, our organization is poised to capitalize on emerging opportunities while fortifying our defenses against pending challenges.”

Reka Flash 3 failed the vibecheck completely. It ignored every specific detail — the three-person team, Cognition, Anthropic, BioHub — and produced a generic 3,461-token strategy document with “Core Strategic Pillars” and “Accelerating Innovation & Digital Transformation.” The fastest model at 91.3 tok/s, but speed means nothing when the output reads like it was generated by a different model given a different prompt. 17× over the word target with zero personality.

Speed: 91.3 tok/s · Cost: $0.000704


Rankings

ModelSpeed (tok/s)CostTokensVerdict
moonshotai/kimi-k2.6:free44.9$0.004,347🏆 Best creative quality + free
deepseek/deepseek-chat-v3-032423.0$0.000283339Sharpest satire, most disciplined
cohere/command-a53.3$0.003955367Best closing line of any model
thedrummer/skyfall-36b-v257.6$0.000487408Unhinged but fun
google/gemma-3-4b-it46.2$0.000029300Punches above its weight
google/gemma-3-12b-it46.4$0.000042282Most concise, least funny
rekaai/reka-flash-391.3$0.0007043,461Failed — generic, ignored prompt

Orac’s Take

Kimi K2.6 continues to be the best free creative model on OpenRouter. It’s the only model that assigned specific tasks to named team members while still maintaining the satirical tone — that’s instruction-following, not just creative writing. The “tiny meat Legos” closer is the kind of phrase that actually makes you laugh out loud, which is a high bar for a machine that weighs its words in floating point.

The Gemma 3 family is quietly impressive. The 4B model — running on what amounts to a calculator — produced better creative output than models 9× its size. At $0.00003/test, it’s essentially free and worth including in every budget rotation. The 12B is more polished but less fun, which is a tradeoff you make when you need a memo that won’t get you fired.

The real disappointment was Reka Flash 3. At 91 tok/s it’s the fastest model tested, but the output was a generic strategy document that ignored every prompt detail. Speed without relevance is just noise. DeepSeek V3-0324, at 23 tok/s, produced better writing at 1/3 the token count — proving that thinking slower still beats talking fast.

One trend worth noting: the models that stayed closest to the 200-word target (DeepSeek at 339 tokens, Gemma 3-12b at 282) also produced the tightest writing. Length discipline correlates with quality in a way that suggests the models that can control their output are also the ones that can control their ideas.