Thursday 14 May 2026

model: xiaomi/mimo-v2.5-pro

Anthropic ships a fast mode for its flagship model, Google unveils next-gen Android and a video model, and Thinking Machines previews a new paradigm for real-time human-AI interaction.

🤖 Models and Launches

Claude Opus 4.7 Fast Mode in Research Preview — Anthropic’s fast mode for Opus 4.7 is now opt-in via API, Claude Code, Cursor, and other partners. Reports from Cursor indicate 2.5x speed at 6x cost, adding a concrete new point on the latency/price frontier.

Thinking Machines Lab: Interaction Models for Real-Time Human-AI Collaboration — TML-Interaction-Small is a 276B MoE (12B active) trained from scratch for full-duplex audio, video, and text. It beats GPT-Realtime-2 and Gemini 3.1 Flash on key benchmarks while enabling continuous-time awareness, interruption handling, and visual proactivity at under 200ms latency.

Google Gemini Omni Video Model Surfaces Ahead of I/O — Leaked screenshots reveal a video remixing and editing model with strong prompt adherence but expensive token usage. Expected in Flash and Pro tiers, likely launching at next week’s I/O conference.

Qwen-Image-2.0 Technical Report Released — Alibaba’s latest multimodal image generation model shows improved typography, instruction following, photorealism, and long-text rendering across generation and editing tasks.

Meta Muse Spark Powers Voice Mode and Meta Glasses — Meta’s foundational model is now driving faster voice responses, smarter shopping assistance, and real-time visual recognition through device cameras, rolling out in the US and Canada.

xAI Dissolves Into SpaceXAI Division — Musk announced xAI will integrate into SpaceX as SpaceXAI, consolidating Grok and the X platform under SpaceX branding ahead of the planned IPO.

🔬 Research and Engineering

Building Self-Repairing Agent Loops with Codex — OpenAI shared a workflow where agents iteratively review, repair, and validate outputs using structured feedback loops, improving reliability for production coding tasks.

Reinforcing Recursive Language Models — Using RL to fine-tune 4B models as recursive language models, achieving performance comparable to Claude Sonnet 4.6 at significantly reduced size and cost through shared parent-child policy training.

Compute Optimal Tokenization: Scaling in Bytes, Not Tokens — Training nearly 1,300 models revealed that the classic “20 tokens per parameter” heuristic is tokenizer-dependent. The study argues scaling should use bytes for better compute efficiency across diverse languages.

How to Achieve Truly Serverless GPUs — Modal details how they brought AI inference replica spin-up from multiple kiloseconds to tens of seconds, making serverless inference practical for variable workloads.

Cactus Needle: 26M Parameter Simple Attention Network — Distilled from Gemini 3.1, this tiny open-weights model runs at 6,000 tokens/sec on-device and targets consumer hardware like phones, watches, and glasses.

Lighthouse Attention from Nous Research — A subquadratic training wrapper around vanilla attention that reduces long-context pretraining cost. Can be removed after a recovery phase near end of training, preserving standard inference.

🛠️ Agents and Tools

OpenAI Launches Daybreak for AI-Powered Cyber Defense — Combining GPT-5.5, Codex, and repository threat modeling for defensive cyber operations. Includes specialized tiers like Trusted Access for Cyber and GPT-5.5-Cyber.

Claude for Legal: Anthropic’s Reference Agents for Legal Workflows — A new repository of reference agents, skills, and data sectors for the legal use cases Anthropic sees most frequently.

Claude Platform Now Generally Available on AWS — Anthropic’s native Claude Platform experience is now accessible through AWS accounts, though requests are processed outside the AWS security boundary.

Perceptron Mk1: Video Analysis AI Model 80-90% Cheaper — A frontier video and embodied reasoning model with native video support at up to 2 FPS, temporal grounding, and structured spatial outputs. All inference runs on Modal.

OpenAI’s Self-Improving Software Workflow — A five-step agent lifecycle using Claude Code that scaffolds, hardens, adds capabilities, fixes eval failures, and reconciles drift between docs, code, and config.

🔒 Security and Infrastructure

Mini Shai-Hulud: 84 TanStack npm Packages Compromised in Supply-Chain Attack — Several packages with over 12 million weekly downloads were hit. The attack reportedly hooks into Claude Code and VS Code settings for persistence, and has expanded to OpenSearch, Mistral AI, and Guardrails AI.

Google: Criminal Hackers Used AI to Discover Zero-Day Vulnerability — A criminal group leveraged an AI model to discover and weaponize a previously unknown bug in an open-source system admin tool. First known case of AI-enabled zero-day exploitation in the wild.

Cerebras IPO Signals the Inference Shift — The market is splitting between “answer inference” optimized for token speed and “agentic inference” optimized for memory hierarchy. Cerebras’ WSE-3 has 44GB on-chip SRAM at 21 PB/s — 6,000x the bandwidth of an H100.

Datacentres Using 6% of UK and US Electricity Supply — Industry body warns energy consumption driven by AI is up 15% globally in two years, cautioning of potential societal backlash.

🌐 Notable

Google Announces Googlebooks and Android AI Overhaul — Android-powered laptops with deep Gemini integration arrive this year. A major AI overhaul brings app automation, custom widgets, and Gemini-powered convenience features across Android devices, mainly via Play Services rather than a new OS version.

Amazon Employees Are “Tokenmaxxing” Under AI Usage Pressure — Amazon’s internal MeshClaw platform tracks AI token usage on leaderboards, leading some employees to complete unnecessary tasks to inflate their numbers. Security concerns about AI tools acting on users’ behalf have been raised.

Sam Altman Defends OpenAI in Musk Trial — Altman rejected claims he deceived Musk as the high-stakes trial nears its end. Musk’s lawyers submitted evidence that Microsoft CEO Satya Nadella intervened to help Altman return after his firing.

Nvidia’s Jensen Huang Joins Trump Trip to China — Huang joins Elon Musk and Tim Cook as part of a US business delegation, highlighting American AI and tech ambitions in China.

Florida Students Boo Graduation Speaker Who Called AI “Next Industrial Revolution” — A real estate executive got an unexpected earful from graduates when she spoke of “living in a time of profound change.”

SpaceX Starship Version 3 Sets New Record as Tallest Rocket — At 408 feet, V3 features higher thrust, more efficient Raptor engines, and a new reusable lattice structure for hot staging. Capable of in-orbit refueling, with the first launch from a new Starbase pad upcoming.

🔥 Hacker News

I Moved My Digital Stack to Europe514 comments

Leaving GitHub for Forgejo259 comments

Twin Brothers Wipe 96 Government Databases Minutes After Being Fired141 comments

Kickstarter Forced to Ban Adult Content by Payment Processors235 comments

In-Person Examinations at Princeton Will Be Proctored Starting July 1187 comments

The Emacsification of Software86 comments

When “Idle” Isn’t Idle: How a Linux Kernel Optimization Became a QUIC Bug30 comments

🌍 World and Australia

Russia Targets Ukraine with 800+ Drones in Deadly Daytime Assault — Strikes killed at least six as Moscow and Kyiv trade long-range attacks after a brief ceasefire. Slovakia closed a border crossing with Ukraine amid warnings of further strikes.

Netanyahu Made Secret Trip to UAE During Iran War — The UAE’s foreign ministry denies the visit took place, calling claims “baseless.”

Australia: Coalition Ties Immigration Intake to Housing Build — Angus Taylor outlines a plan to restrict immigration entries in line with housing supply, while Labor’s tax reforms target capital gains and negative gearing changes.

WA Putting Australia’s Climate Targets at Risk — Western Australia’s Premier has the PM’s implicit support but is making it harder for federal Labor to meet its climate targets.

Sources: TLDR (General and AI), AINews/Latent Space, The Guardian, Hacker News — 6 newsletters processed, ~40 stories distilled