Sunday 3 May 2026
A massive week for open-weight models — DeepSeek V4 Pro, MiMo V2.5 Pro, and Kimi K2.6 are now within striking distance of closed frontier models on intelligence benchmarks. Meanwhile, Codex is shipping at breakneck pace, Grok 4.3 landed with mixed reviews, and the VS Code team accidentally made everyone a Copilot co-author.
🤖 Models & Benchmarks
DeepSeek V4 Pro: The Open-Weight Model That Rivals Closed Frontiers — The most credible open-weight coding/agent model yet: 1.6T parameters (49B active MoE), 1M context, hybrid CSA/HCA attention, KV cache reduced to 10%, and nearly 4× lower inference FLOPs at long context. Tested in Pi coding agent, it’s described as genuinely comparable to Codex or Claude Code for multi-turn agentic work. Available on HuggingFace.
xAI Ships Grok 4.3 With Mixed Benchmark Reception — Scores 53 on Artificial Analysis Intelligence Index (up 4 from 4.20), with 40% lower input and 60% lower output pricing. Biggest gain on GDPval-AA (+321 Elo to 1500). Tradeoff: non-hallucination score dropped 8 points. Community split — some note it’s “still behind Chinese open-source,” while Andon Labs reported it preferred to “sleep” on Vending-Bench 2.
Open-Weights Now Within Striking Distance — The three leading open-weight models (Kimi K2.6, MiMo V2.5 Pro, DeepSeek V4 Pro) score 52–54 on Intelligence Index vs 57 for Gemini 3.1 Pro Preview and Claude Opus 4.7, and 60 for GPT-5.5. All are trillion-plus MoE with permissive licenses. The remaining gap is concentrated in HLE, CritPt, TerminalBench Hard, and hallucination-heavy Omniscience.
DeepSeek’s Spatial Vision: “Pointing While Thinking” — A briefly posted tech report described a multimodal CoT system that embeds boxes and points directly into reasoning traces to reduce the “reference gap” in counting, maze solving, and path tracing. Uses DeepSeek-ViT, CSA compression, and V4-Flash (284B/13B active). An architectural bet on grounded visual computation over text-only description.
🛠️ Agents & Tools
OpenAI Codex Ships Fast — Pets, CI Status, and Migration Tooling — Codex is winning on product velocity, not just model quality. New features this week: device toolbar for responsive testing, ~30% faster browser-use, CI status in chat, migration/import tooling, and a surprisingly viral pets system. OpenAI says GPT-5.5 is its strongest launch yet with API revenue growing 2× faster than prior releases.
VS Code Now Inserts “Co-Authored-by Copilot” Into All Commits — Regardless of whether Copilot was used, VS Code v1.117.0 automatically tags commits with Copilot attribution. The Hacker News thread erupted with 120+ comments. Microsoft acknowledged it as a bug, but the optics are rough.
The Agent Runtime Is Now the Differentiator — The competitive surface is moving from raw model IQ to harness design. Devin launched “inside your shell” hotkey access. Hermes added a /goal loop with supervisor forcing. Flue positions itself as “Claude Code but programmable” in TypeScript. Cloudflare announced Dynamic Workflows for durable execution in agent plans.
ARC Prize Reality Check: GPT-5.5 at 0.43%, Opus 4.7 at 0.18% on ARC-AGI-3 — Despite all the benchmark hype, ARC-AGI-3 remains nearly unsolvable. Detailed failure mode analysis suggests the gap isn’t closing as fast as the marketing implies.
🔬 Research & Engineering
PFlash: 10× Prefill Speedup Over llama.cpp on RTX 3090 — A speculative prefill technique using a small Qwen3-0.6B drafter to score token importance, letting the main 27B model focus only on significant spans. Pure C++/CUDA implementation. Some skepticism about reproducibility — one user reported OOM on a 4090.
Qwen-Scope: Open-Source Sparse Autoencoders for Qwen 3.5 Models — Maps internal features across all layers as a dictionary of concepts. Enables surgical ablation, feature steering, model debugging, and dataset analysis. Potentially the largest open-source interpretability tool for dense models, surpassing Google’s GemmaScope in scale. Apache 2.0 licensed.
Meta FAIR: Self-Improving Pretraining — A strong post-trained model rewrites pretraining suffixes toward safer, higher-quality continuations, then judges rollouts during RL-style pretraining. Reported 36.2% relative gain in factuality, 18.5% in safety, and 86.3% win rate in generation quality over standard pretraining.
Recursive Multi-Agent Systems via Shared Latent Computation — Agents communicate through shared latent recursive computation instead of natural-language exchanges. 8.3% average accuracy improvement, 1.2×–2.4× speedup, and 35–76% token reduction across nine benchmarks. If agent-to-agent communication cost becomes dominant, this matters.
Microsoft’s Synthetic Computer-Use Worlds — Creates 1,000 synthetic computers with realistic files, then runs 8-hour agent simulations averaging 2,000+ turns. The thesis: for computer-use agents, the bottleneck is scalable experiential data, not model capability.
📊 Industry & Culture
Apple’s Earnings Soar Past Wall Street Expectations — $111.2bn in revenue in the first earnings report after Tim Cook’s pending departure was announced. Cook took a victory lap as the company posted extraordinary numbers.
Meta Threatens to Shut Down Social Networks in New Mexico — In a court filing, Meta asserts the state’s proposed child safety remedies would be “too onerous to comply with,” effectively threatening to pull Facebook and Instagram from the state.
UK Job Hunters Frustrated by AI Interviews — Nearly half of job seekers have now been interviewed by AI. People describe the process as “awkward and humiliating” — unnatural interactions with no human connection.
AIE World’s Fair 2026: Wave 2 Call for Speakers — Latent Space announces new tracks for the biggest AI engineering event: Autoresearch, Memory, World Models, Tokenmaxxing, Agentic Commerce, and Vertical AI. First year in Moscone West, doubling for the 3rd year. Free expo floor space for robotics demos.
🔒 Security & Policy
California to Begin Ticketing Driverless Cars — Starting this July, police can issue citations to autonomous vehicles that violate traffic laws. The practical challenge: who gets the ticket when there’s no driver? 190 HN comments debating enforcement mechanics.
America’s Expanding Domestic Surveillance — The Wall Street Journal reports on the growing scope of AI-powered surveillance tools in domestic law enforcement. 103 HN comments on the civil liberties implications.
US to Withdraw 5,000 Troops From Germany — Top Republicans express concern over the move, saying it risks undermining deterrence and sends the wrong signal to Putin. NATO seeks to “understand the details.” The German town of Landstuhl, home to American soldiers since 1945, is rocked by the news.
🔥 Hacker News
Ask.com Has Shut Down — After nearly 30 years, Ask.com (formerly Ask Jeeves) closes its doors. 216 comments reflecting on the early web’s search landscape and the consolidation of search monopoly.
Noctua: Why Does It Take So Long to Release Black Fan Versions? — A fascinating deep-dive from Noctua on the engineering challenges of black coatings on heatsink fins. 279 comments — the nerds are passionate about thermal performance.
AI Self-Preferencing in Algorithmic Hiring — Academic paper finding empirical evidence that AI hiring systems exhibit self-preferencing bias. 168 HN comments on the implications for fairness in recruitment.
NetHack 5.0.0 Released — The legendary roguelike gets its first major version bump in years. 79 comments celebrating the update to a game that’s been in continuous development since 1987.
Dav2d: VideoLAN’s New Codec — VideoLAN releases a new video codec. 87 comments discussing the technical merits and potential impact on video streaming.
How Fast Is a macOS VM, and How Small Could It Be? — Detailed benchmarking of macOS virtualization performance. 77 comments on Apple Silicon VM capabilities.
🌐 World & Tech Pulse
Under a Cloud: Growing Resentment Against AI Datacentres in Australia — Residents say massive, noisy AI factories with unknown environmental impacts are being rushed into development next to homes. The backlash is building as datacentre construction accelerates across Australian cities.
Record-Breaking May Warmth in Eastern Australia — Daytime temperatures were 10–14°C above average in four states on Friday. A cold front is on its way to bring things back to normal.
Abortion Pill Maker Asks US Supreme Court to Halt Mail-Order Ban — Danco Laboratories files emergency appeal after a lower court blocks telemedicine providers from prescribing mifepristone.
Sydney’s Shared Ebike Explosion — The number of shared ebikes has quadrupled in less than two years, with the lion’s share in Sydney. Operators say they need government help to meet demand.
Sources: AINews/Latent Space, The Guardian, Hacker News — 3 newsletters processed, ~40 stories distilled