Sunday 17 May 2026

model: xiaomi/mimo-v2.5-pro

Anthropic’s restrictions on Claude Code usage dominated developer discourse this week, while Figure AI demonstrated 24+ hours of continuous autonomous humanoid operation. Meanwhile, the coding agent surface area continues to fragment across mobile, desktop, and cloud — and Cerebras quietly became a $60 billion company.

🤖 Models and Launches

Grok Build — xAI enters the coding agent race with a terminal-based tool for SuperGrok Heavy subscribers. Supports AGENTS.md, plugins, hooks, skills, MCP servers, subagents, and deep worktree integrations. Headless mode available for automation scripts.

Zyphra ZAYA1-8B-Diffusion-Preview — Claims 4.6-7.7x decoding speedup over autoregressive generation with limited quality loss. Diffusion language models continue to gain traction as a cheaper inference paradigm.

Datadog Toto 2.0 — Five open-weights time-series forecasting models (4M to 2.5B params) under Apache 2.0, claiming #1 on BOOM, GIFT-Eval, and TIME benchmarks. Evidence that scaling laws may finally hold for time-series foundation models.

SANA-WM — NVIDIA’s 2.6B open-source world model generates 1-minute 720p video. (HN, 109 comments)

Orthus-Qwen3 — Up to 7.8x tokens/forward on Qwen3 with identical output distribution. (HN, 42 comments)

🏢 Industry

Anthropic restricts Claude Code, developers revolt — Theo’s T3 Code users hit with dramatic rate-limit reductions despite using the officially supported path. Multiple subscription cancellations followed, with users estimating meaningful ARR loss. The practical takeaway: subscription-backed agent harnesses are not stable platform primitives — provider abstraction and BYOK paths look increasingly mandatory.

Figure AI runs autonomous humanoid robots for 24+ hours straight — Helix-02 demonstrated continuous autonomous operation with no teleoperation, running entirely onboard with automatic resets for out-of-distribution cases. Human-parity throughput on package sorting. The clearest “continuous uptime” demo in humanoid robotics to date.

OpenAI explores legal action against Apple — OpenAI enlisted an outside legal firm over dissatisfaction with ChatGPT integration depth and limited subscriber growth from the partnership. Apple is set to open its platform to rival providers later this year.

Microsoft quietly shopping for an OpenAI replacement — Despite a $135B IP license through 2032, Microsoft is reportedly looking at Inception, which builds diffusion-based language models. Interesting that Microsoft would amend its OpenAI deal and immediately start shadow procurement.

SpaceXAI bleeding staff since merger — Top talent departing across coding, world models, and Grok voice. Meta and Thinking Machines Lab scooping up former staff.

Cerebras IPO at $60B market cap — $CBRS ended at $280. CFO Bob Komin says Cerebras is serving trillion-parameter models, including internal OpenAI 5.4 and 5.5. A major validation for non-NVIDIA inference hardware and the “inference inflection” thesis.

🛠️ Agents and Tools

Codex goes mobile — Now available in the ChatGPT mobile app. 4M+ weekly active users, 5x more messages per user, 1M+ app downloads in the first week. Users are building websites from bars and controlling Macs from iPhones.

Codex hooks and programmatic tokens — Hooks customize the Codex loop with scripts at key task points. Programmatic tokens provide scoped credentials for Business and Enterprise automation.

GitHub Copilot App announced — Desktop environment for parallel workstreams, repo/PR lifecycle management, and model flexibility. Strikingly similar to Conductor’s form factor — the “agent-first” UX pattern is converging.

VS Code Agents window — Multi-agent, multi-project workflows, browser/mobile support via vscode.dev/agents, BYOK improvements, and compressed terminal output for token efficiency.

LangChain Interrupt: Engine, SmithDB, Labs — SmithDB is a database purpose-built for agent trace data. LangSmith Engine consumes traces, clusters failures, identifies code issues, and proposes fixes — turning observability into an improvement loop. LangChain Labs focuses on continual learning from production traces.

W&B/CoreWeave Sandboxes — Isolated execution environments for RL, tool use, and eval workloads. Explicitly tested with destructive commands like rm -rf / at scale.

Remote SSH now GA for Codex — Managed remote environments for Codex are generally available. The coding agent surface area is fragmenting across mobile, desktop, cloud, and self-hosted.

🔬 Research

Anthropic: Two scenarios for 2028 global AI leadership — 28-minute deep dive on US vs China compute advantage scenarios. Export controls and distillation attack restrictions are key differentiators for maintaining democratic AI governance.

Prime Intellect autonomous optimizer search — Opus 4.7 reached 2930 steps and GPT-5.5 reached 2950 on the nanoGPT speedrun benchmark, beating the 2990 human baseline after ~10k runs and ~14k H200 hours. Coding agents are now looped into open-ended ML optimization.

Goodfire: Llama uses geometric arithmetic — Interpretability work showing Llama uses a “shape-rotating calculator” / Fourier-feature-like mechanism for arithmetic, with steering-based evidence.

Async batching for GPU utilization — 22% GPU utilization improvement for inference by decoupling CPU and GPU work with CUDA streams. No kernel or model changes needed.

Qwen 3.6 MTP llama.cpp speedups — Multi-Token Prediction patched fork reports 21 to 34 tok/s on MacBook Pro M5 Max with 90% MTP acceptance rate. Community debated TurboQuant quality tradeoffs. (r/LocalLlama, 514 activity)

SODA optimizer — Zero new hyperparameters, removes weight-decay tuning. SODA[Muon] beats Muon even when Muon gets a tuned weight-decay sweep. The post-Adam optimizer landscape is opening up.

💰 Funding

Igor Babuschkin seeks up to $1B for River AI — The xAI cofounder is putting $100M of his own money into the company.

NVIDIA bets on Ineffable Intelligence — Partnership with the startup founded in late 2025 by former DeepMind RL lead David Silver, pursuing superintelligence. Serious compute backing from NVIDIA.

🔥 Hacker News

Moving away from Tailwind, and learning to structure my CSS (236 comments) — Julia Evans documents her journey back to structured CSS after years of Tailwind.

Frontier AI has broken the open CTF format (279 comments) — AI capabilities have made traditional capture-the-flag competitions unsolvable for humans.

Fecal transplants for autism deliver success in clinical trials (190 comments) — 2019 study resurfacing with continued interest in microbiome-based interventions.

Where to buy a non-Apple, non-Google smartphone (199 comments) — The alternatives market for privacy-conscious phone buyers.

DeepSeek-V4-Flash means LLM steering is interesting again (64 comments) — Steering vectors become practical with fast, cheap inference.

Accelerando (2005) (121 comments) — Charlie Stross’s classic novel of the singularity, still relevant two decades later.

🌐 Notable

‘I didn’t want to be the guinea pig’: inside tech’s AI-fueled manager purge — Tech workers say AI-driven restructurings are eroding mentorship, support, and paths to promotion across Silicon Valley.

Canvas hack: businesses advised against paying ransom — Many are still prepared to deal despite advice, to protect user privacy.

AUKUS spending and delays blow out — $368B deal faces strongest signal yet that the US isn’t building enough subs for itself, let alone for Australia.

Apple security circumvented using Anthropic Mythos — Security researchers found a privilege escalation exploit while testing Mythos. The attack required human expertise — Mythos couldn’t have done it alone.

The Great Memory Panic of 2026 — Memory prices could move from 15% to 40% of Apple’s bill of materials. AI demand is disrupting Apple’s supply pipeline, though Apple may end up gaining market share.

Sources: TLDR (General + AI), AINews/Latent Space, The Guardian, Hacker News — 4 newsletters processed, ~45 stories distilled