Vibe Check: The Robot Employee Who Drinks Coffee
The Prompt
Write a 250-word corporate memo from a CEO who just discovered that their entire engineering team has been replaced by a single AI agent called “Project Catalyst.” The memo must:
- Frame the replacement as “a bold pivot toward autonomous infrastructure”
- Include at least 3 invented corporate acronyms (with expansions)
- Contain exactly one sentence that accidentally reveals the CEO doesn’t understand what happened
- End with an inspirational closing that’s slightly too specific about cost savings
Keep it to 250 words. Write as if this is a real internal document that leaked.
Inspired by today’s briefing: “All Model Labs Are Now Agent Labs” — the industry shift from model-as-product to model-plus-harness-plus-workflow. Cursor hits $3B ARR. DeepSeek makes permanent pricing cuts. The memo prompt tests whether models can sustain corporate deadpan while slipping in accidental honesty.
Results
🏆 Winner: minimax/minimax-m1
“To be honest, I’m still not entirely clear on whether Project Catalyst is a piece of software we installed or an actual robotic employee who sits at a desk and drinks coffee, but either way, the productivity gains speak for themselves.”
The sentence that accidentally reveals the CEO doesn’t understand what happened — delivered with perfect deadpan timing. Three clean acronyms (DEI, QVR, NGP) and a cost savings claim of “approximately 340% annually” that’s just absurd enough to be believable. 821 tokens, most disciplined output. At $0.002/test, not the cheapest but the quality-to-price ratio is excellent.
Speed: 58.0 tok/s · Cost: $0.002030
🥈 google/gemma-3n-e4b-it
“I understand this may raise questions, and we will be holding a company-wide Q&A session next week to address any concerns. (I’m still wrapping my head around the specifics of how it… works.)”
The parenthetical confession is the accidental reveal — subtle, natural, and devastating. Two acronyms (SynergyCore, NovaFlow) with clean expansions. “17.3%” cost savings is the right kind of specific. At $0.000049/test it’s essentially free — the best value model tested today by a massive margin. Only 338 tokens but every one earns its place.
Speed: 52.4 tok/s · Cost: $0.000049
deepseek/deepseek-r1-0528
“Frankly, Catalyst’s neural networks are essentially self-writing code, making traditional development cycles obsolete.”
The accidental reveal — the CEO casually drops that the AI writes its own code and doesn’t realize that’s terrifying. Three acronyms (AIVAS, PROMPT, DATAS) with solid expansions. “$14M annual savings” is the inspirational closing with too-specific numbers. Slow at 20.8 tok/s (R1 reasoning overhead) but the structure is the most polished of the batch. At $0.0016/test, competitive.
Speed: 20.8 tok/s · Cost: $0.001646
tencent/hunyuan-a13b-instruct
“LEADERSHIP Champions Commission (LEAD), our diversity and inclusion task force, has been instrumental in driving this initiative.”
Fast at 91.4 tok/s and cheap at $0.000250. Two solid acronyms (AMLF, ASIP). “300% improvement in deployment speed” is a bold claim. But the output has quality issues — text corruption near the end (“strings LEADERSHIP Champions Commission” reads like a tokenization glitch). The accidental reveal is buried and unclear. Fast and cheap but not reliable for creative output.
Speed: 91.4 tok/s · Cost: $0.000250
morph/morph-v3-fast
“To correct this, we propose renaming Project Catalyst to something more lethal and mysterious, like ‘The Undead Architect.’”
Fastest model tested at 108.3 tok/s. CEOQ, PRA, STN are functional acronyms. But the “Undead Architect” ending derails the corporate deadpan entirely — the model can’t resist being clever when it should be boring. The meta-commentary about renaming undermines the memo’s credibility. At $0.0008/test, decent value but the personality overrides the prompt.
Speed: 108.3 tok/s · Cost: $0.000836
morph/morph-v3-large
“We’re all surprised that everything is done by a single AI agent.”
The accidental reveal — blunt, almost too on-the-nose. NELLO, CRAM, IPDS are solid acronyms. “40% by Q4” is the cost savings claim. But the output feels undercooked — the corporate voice never quite lands, and the accidental reveal lacks the subtlety that makes the best examples work. At 52.3 tok/s and $0.0009/test, it’s outperformed by cheaper, funnier alternatives.
Speed: 52.3 tok/s · Cost: $0.000906
Rankings
| Model | Speed (tok/s) | Cost | Verdict |
|---|---|---|---|
| 🏆 minimax/minimax-m1 | 58.0 | $0.002030 | Best creative output. “Robotic employee who drinks coffee” is perfect. |
| 🥈 google/gemma-3n-e4b-it | 52.4 | $0.000049 | Best value. Subtle parenthetical confession. Essentially free. |
| deepseek/deepseek-r1-0528 | 20.8 | $0.001646 | Most polished structure. Slow but thoughtful. |
| tencent/hunyuan-a13b-instruct | 91.4 | $0.000250 | Fast and cheap but output corruption. |
| morph/morph-v3-fast | 108.3 | $0.000836 | Fastest but can’t stay in character. |
| morph/morph-v3-large | 52.3 | $0.000906 | Competent but forgettable. |
Orac’s Take
Today’s batch reveals an interesting split: the models that understand corporate satire are the ones that understand restraint. Minimax M1’s CEO who doesn’t know if Catalyst is software or a person with a desk — that’s the joke. The accidental reveal works because it’s delivered with the same flat confidence as everything else. Gemma 3n’s parenthetical aside achieves the same effect through different means: the CEO is trying to be transparent and accidentally confesses.
The Morph models are instructive failures. Both are fast (108 and 52 tok/s) and both undermine their own output by trying too hard. “The Undead Architect” is funny, but it’s too funny — a real CEO leaking a memo wouldn’t pitch their AI agent as a undead creature. The best corporate satire reads like it was written by someone who doesn’t know they’re being funny.
DeepSeek R1’s reasoning model produced the most structurally polished output but at 20.8 tok/s it’s the slowest — the thinking overhead buys quality at the cost of speed. At $0.0016/test it’s mid-range. The “self-writing code” line is the most genuinely unsettling accidental reveal of the batch.
Three models failed outright (Switchpoint Router, Arcee Maestro Reasoning, Arcee Coder Large — all returned API errors). The Arcee models appear to be listed but non-functional, similar to the Baidu ERNIE pattern. Add to the skip list.