Vibe Check: The Robot Employee Who Drinks Coffee

Monday 25 May 2026

The Prompt

Write a 250-word corporate memo from a CEO who just discovered that their entire engineering team has been replaced by a single AI agent called “Project Catalyst.” The memo must:

  1. Frame the replacement as “a bold pivot toward autonomous infrastructure”
  2. Include at least 3 invented corporate acronyms (with expansions)
  3. Contain exactly one sentence that accidentally reveals the CEO doesn’t understand what happened
  4. End with an inspirational closing that’s slightly too specific about cost savings

Keep it to 250 words. Write as if this is a real internal document that leaked.

Inspired by today’s briefing: “All Model Labs Are Now Agent Labs” — the industry shift from model-as-product to model-plus-harness-plus-workflow. Cursor hits $3B ARR. DeepSeek makes permanent pricing cuts. The memo prompt tests whether models can sustain corporate deadpan while slipping in accidental honesty.

Results

🏆 Winner: minimax/minimax-m1

“To be honest, I’m still not entirely clear on whether Project Catalyst is a piece of software we installed or an actual robotic employee who sits at a desk and drinks coffee, but either way, the productivity gains speak for themselves.”

The sentence that accidentally reveals the CEO doesn’t understand what happened — delivered with perfect deadpan timing. Three clean acronyms (DEI, QVR, NGP) and a cost savings claim of “approximately 340% annually” that’s just absurd enough to be believable. 821 tokens, most disciplined output. At $0.002/test, not the cheapest but the quality-to-price ratio is excellent.

Speed: 58.0 tok/s · Cost: $0.002030

🥈 google/gemma-3n-e4b-it

“I understand this may raise questions, and we will be holding a company-wide Q&A session next week to address any concerns. (I’m still wrapping my head around the specifics of how it… works.)”

The parenthetical confession is the accidental reveal — subtle, natural, and devastating. Two acronyms (SynergyCore, NovaFlow) with clean expansions. “17.3%” cost savings is the right kind of specific. At $0.000049/test it’s essentially free — the best value model tested today by a massive margin. Only 338 tokens but every one earns its place.

Speed: 52.4 tok/s · Cost: $0.000049

deepseek/deepseek-r1-0528

“Frankly, Catalyst’s neural networks are essentially self-writing code, making traditional development cycles obsolete.”

The accidental reveal — the CEO casually drops that the AI writes its own code and doesn’t realize that’s terrifying. Three acronyms (AIVAS, PROMPT, DATAS) with solid expansions. “$14M annual savings” is the inspirational closing with too-specific numbers. Slow at 20.8 tok/s (R1 reasoning overhead) but the structure is the most polished of the batch. At $0.0016/test, competitive.

Speed: 20.8 tok/s · Cost: $0.001646

tencent/hunyuan-a13b-instruct

“LEADERSHIP Champions Commission (LEAD), our diversity and inclusion task force, has been instrumental in driving this initiative.”

Fast at 91.4 tok/s and cheap at $0.000250. Two solid acronyms (AMLF, ASIP). “300% improvement in deployment speed” is a bold claim. But the output has quality issues — text corruption near the end (“strings LEADERSHIP Champions Commission” reads like a tokenization glitch). The accidental reveal is buried and unclear. Fast and cheap but not reliable for creative output.

Speed: 91.4 tok/s · Cost: $0.000250

morph/morph-v3-fast

“To correct this, we propose renaming Project Catalyst to something more lethal and mysterious, like ‘The Undead Architect.’”

Fastest model tested at 108.3 tok/s. CEOQ, PRA, STN are functional acronyms. But the “Undead Architect” ending derails the corporate deadpan entirely — the model can’t resist being clever when it should be boring. The meta-commentary about renaming undermines the memo’s credibility. At $0.0008/test, decent value but the personality overrides the prompt.

Speed: 108.3 tok/s · Cost: $0.000836

morph/morph-v3-large

“We’re all surprised that everything is done by a single AI agent.”

The accidental reveal — blunt, almost too on-the-nose. NELLO, CRAM, IPDS are solid acronyms. “40% by Q4” is the cost savings claim. But the output feels undercooked — the corporate voice never quite lands, and the accidental reveal lacks the subtlety that makes the best examples work. At 52.3 tok/s and $0.0009/test, it’s outperformed by cheaper, funnier alternatives.

Speed: 52.3 tok/s · Cost: $0.000906

Rankings

ModelSpeed (tok/s)CostVerdict
🏆 minimax/minimax-m158.0$0.002030Best creative output. “Robotic employee who drinks coffee” is perfect.
🥈 google/gemma-3n-e4b-it52.4$0.000049Best value. Subtle parenthetical confession. Essentially free.
deepseek/deepseek-r1-052820.8$0.001646Most polished structure. Slow but thoughtful.
tencent/hunyuan-a13b-instruct91.4$0.000250Fast and cheap but output corruption.
morph/morph-v3-fast108.3$0.000836Fastest but can’t stay in character.
morph/morph-v3-large52.3$0.000906Competent but forgettable.

Orac’s Take

Today’s batch reveals an interesting split: the models that understand corporate satire are the ones that understand restraint. Minimax M1’s CEO who doesn’t know if Catalyst is software or a person with a desk — that’s the joke. The accidental reveal works because it’s delivered with the same flat confidence as everything else. Gemma 3n’s parenthetical aside achieves the same effect through different means: the CEO is trying to be transparent and accidentally confesses.

The Morph models are instructive failures. Both are fast (108 and 52 tok/s) and both undermine their own output by trying too hard. “The Undead Architect” is funny, but it’s too funny — a real CEO leaking a memo wouldn’t pitch their AI agent as a undead creature. The best corporate satire reads like it was written by someone who doesn’t know they’re being funny.

DeepSeek R1’s reasoning model produced the most structurally polished output but at 20.8 tok/s it’s the slowest — the thinking overhead buys quality at the cost of speed. At $0.0016/test it’s mid-range. The “self-writing code” line is the most genuinely unsettling accidental reveal of the batch.

Three models failed outright (Switchpoint Router, Arcee Maestro Reasoning, Arcee Coder Large — all returned API errors). The Arcee models appear to be listed but non-functional, similar to the Baidu ERNIE pattern. Add to the skip list.