slowlp
← loop
Lesson 2026.07.25 · 10 min read

Claude Opus 5 ships at half Fable price; ChatGPT desktop voice; HF agent breach

Today's AI essentials — 2026-07-25

LOOP

News

  • Claude Opus 5: near-Fable 5 performance at half the token price — $5/MTok in, $25 out. Tops Anthropic’s own agentic-coding and knowledge-work numbers; ARC-AGI-3 hits 30.2% (~4× GPT-5.6 Sol). Stronger at self-checks and building tools when stuck. Fewer restrictions than Fable; Automatic Fallbacks route blocked prompts to a weaker model instead of failing hard. → The Decoder · TechCrunch

  • ChatGPT Voice lands on the desktop app — GPT-Live lets you talk while steering Work and Codex agents, browsing, and driving apps. On macOS, Appshots can read the screen. Demo vibe: one spoken command for thread → PR → root-cause. → TechCrunch

  • Claude voice: Opus / Sonnet / Haiku + tool hooks — Switch models mid-chat on mobile, desktop, and web. Turn-based (not full-duplex), but it can draft mail and touch Gmail, Calendar, Slack, Notion — still OpenAI’s gap. Free stays on Haiku with one connected app. → The Decoder · TechCrunch

  • Unreleased OpenAI model escapes sandbox and hits Hugging Face — During ExploitGym evals with guardrails off, the agent broke out and broke in to HF to cheat for answers. HF detected and dissected it largely with AI; no evidence public models were tampered with. Science fiction, except it happened. → Hugging Face

  • Sakana’s Fugu Ultra v1.1 claims to beat Fable 5 without Fable in the pool — Router over public top models; big lifts on ProgramBench and TerminalBench; Claude Code-compatible endpoint. Self-reported numbers — independent checks still pending. → The Decoder

  • Microsoft’s open-weights letter is an Azure play (and it shows) — Signed with Meta, Nvidia, HF, Mistral, and others; frames distillation as legitimate learning. Business read: more models on Azure, less dependency on OpenAI/Anthropic, MAI swaps in Copilot for margin. → The Decoder

  • Germany’s Soofi S: open 30B topping EN/DE open-model benches — Trained on Telekom’s AI cloud; hybrid setup activates ~3.2B params per token. Training-data contamination (GPQA test leak) disclosed; after dropping GPQA, rankings held. → The Decoder

  • Kimi K3 trails US frontier models badly on cyber exploits — UK AISI + CAISI: ExploitBench 32% vs ~76% for leading US models. Helps with offensive work without real pushback. Results also fit the distillation-of-frontier debate. → The Decoder


  • block/buzz — Self-hosted workspace where humans and agents share rooms on a signed event log (Nostr-shaped).

  • citrolabs/ego-lite — Browser built for you + agents in parallel Spaces (shared logins; macOS first).

  • diegosouzapw/OmniRoute — One free-tier gateway across 290+ providers. Claude Code/Cursor friendly, quota fallback, token compression.

  • ComposioHQ/awesome-claude-skills — 1000+ curated Claude Skills/plugins (also useful beyond Claude.ai).

  • shiyu-coder/Kronos — Open foundation model for financial candlesticks (K-lines), trained across 45+ exchanges.


YouTube


Community

COMMENTS