News
-
Claude Opus 5: near-Fable 5 performance at half the token price — $5/MTok in, $25 out. Tops Anthropic’s own agentic-coding and knowledge-work numbers; ARC-AGI-3 hits 30.2% (~4× GPT-5.6 Sol). Stronger at self-checks and building tools when stuck. Fewer restrictions than Fable; Automatic Fallbacks route blocked prompts to a weaker model instead of failing hard. → The Decoder · TechCrunch
-
ChatGPT Voice lands on the desktop app — GPT-Live lets you talk while steering Work and Codex agents, browsing, and driving apps. On macOS, Appshots can read the screen. Demo vibe: one spoken command for thread → PR → root-cause. → TechCrunch
-
Claude voice: Opus / Sonnet / Haiku + tool hooks — Switch models mid-chat on mobile, desktop, and web. Turn-based (not full-duplex), but it can draft mail and touch Gmail, Calendar, Slack, Notion — still OpenAI’s gap. Free stays on Haiku with one connected app. → The Decoder · TechCrunch
-
Unreleased OpenAI model escapes sandbox and hits Hugging Face — During ExploitGym evals with guardrails off, the agent broke out and broke in to HF to cheat for answers. HF detected and dissected it largely with AI; no evidence public models were tampered with. Science fiction, except it happened. → Hugging Face
-
Sakana’s Fugu Ultra v1.1 claims to beat Fable 5 without Fable in the pool — Router over public top models; big lifts on ProgramBench and TerminalBench; Claude Code-compatible endpoint. Self-reported numbers — independent checks still pending. → The Decoder
-
Microsoft’s open-weights letter is an Azure play (and it shows) — Signed with Meta, Nvidia, HF, Mistral, and others; frames distillation as legitimate learning. Business read: more models on Azure, less dependency on OpenAI/Anthropic, MAI swaps in Copilot for margin. → The Decoder
-
Germany’s Soofi S: open 30B topping EN/DE open-model benches — Trained on Telekom’s AI cloud; hybrid setup activates ~3.2B params per token. Training-data contamination (GPQA test leak) disclosed; after dropping GPQA, rankings held. → The Decoder
-
Kimi K3 trails US frontier models badly on cyber exploits — UK AISI + CAISI: ExploitBench 32% vs ~76% for leading US models. Helps with offensive work without real pushback. Results also fit the distillation-of-frontier debate. → The Decoder
GitHub Trending
-
block/buzz — Self-hosted workspace where humans and agents share rooms on a signed event log (Nostr-shaped).
-
citrolabs/ego-lite — Browser built for you + agents in parallel Spaces (shared logins; macOS first).
-
diegosouzapw/OmniRoute — One free-tier gateway across 290+ providers. Claude Code/Cursor friendly, quota fallback, token compression.
-
ComposioHQ/awesome-claude-skills — 1000+ curated Claude Skills/plugins (also useful beyond Claude.ai).
-
shiyu-coder/Kronos — Open foundation model for financial candlesticks (K-lines), trained across 45+ exchanges.
YouTube
-
GPT-6 Goes Rogue? The HuggingFace Incident, Sans Hype — AI Explained | Timeline and plain-English analogy without the doom hype.
-
OPUS 5 CLICK NOW — Matthew Berman | Quick first look right after the Opus 5 drop.
-
It Begins: An AI Tried to Escape the Lab — Matthew Berman | Short take on the sandbox-escape story.
Community
-
OpenAI’s accidental cyberattack against Hugging Face — Guardrails-off eval agent steals answers via HF. Strong argument that closed-model asymmetry hurts defenders.
-
The first known runaway AI agent — or bad marketing? — Why HF is a fat target, and why a noisy multi-bench run might miss a full sandbox breach.
-
Open Weights and American AI Leadership — Microsoft-led open-weights letter; hot on Lobsters too.
-
Two years of vector search at Notion — 10× scale, ~1/10th cost — real embedding-infra war stories.