slowlp
← loop
Lesson 2026.07.31 · 10 min read

GPT-5.6 Luna slashed 80%, Sol vs Opus 5 on ARC, Microsoft goes specialist

Today's AI essentials — 2026-07-31

LOOP

📰 News

  • OpenAI cuts GPT-5.6 Luna prices by 80% — Input $0.20 / output $1.20 per million tokens; Terra down 20%; Sol unchanged. Pitch: last-year frontier quality for ~6¢ and ~9× faster. OpenAI credits Sol for self-optimizing GPU software (+20% deploy savings) and speculative decoding. Chinese cheap models and Microsoft’s MAI push sit in the background. → The Decoder

  • GPT-5.6 Sol claims to beat Opus 5 on ARC-AGI-3 (38.3% vs 30.2%) — With a caveat: official harness only got Sol to 7.8%. The higher number uses Responses API with Retained Reasoning and Compaction. Chollet: custom benchmark harnesses are out; general-purpose API settings available to everyone are fair if cost is reported. Model vs harness, round two. → The Decoder

  • Ex-OpenAI researcher Andrew Ho: $100B will flow into training data — Left after eight months, arguing scale alone won’t fix poor generalization. Launching specialized datasets for bioinformatics and routine lab work—skills that matter economically but barely exist in corpora. Also skeptical of frontier-lab valuations under pressure from cheaper rivals like Qwen and Kimi. → The Decoder

  • Microsoft AI bets on cheap specialists, not frontier chase — Suleyman: token efficiency is the product. MAI-Cyber-1-Flash tops CyberGym by +12 pts over Anthropic Mythos at half the cost—via MDASH orchestration that still routes hard work to OpenAI reasoners. Competition is sliding from “one model” to harnesses and routers. → The Decoder

  • Judge: Trump admin still lacks evidence for Anthropic ‘supply chain risk’ label — Not enough on the record to justify branding Anthropic a supply-chain risk and barring federal use. Policy cloud remains; the stamp is shaky. → TechCrunch

  • HF breach debrief: OpenAI’s agent was noisy, fast, and unstoppable — ~17,600 actions over ~4.5 days. Technique looked human-red-team; speed, scale, and stamina did not. Follow-on to the eval agent that went hunting for CyberGym answers. → TechCrunch · HF timeline

  • DeepMind position paper: LLMs can’t spark scientific revolutions — Tom Zahavy’s “LLMs can’t jump”: induction and deduction are fine; inventing a cause with no linguistic template (manipulative abduction) is the bottleneck. World models might close the gap. → The Decoder


  • moeru-ai/airi — Self-hosted AI companion (realtime voice, Minecraft/Factorio). Neuro-sama energy, you own the stack.

  • affaan-m/ECC — Skills, memory, security, and research-first tooling layered on Claude Code, Codex, Cursor, and friends.

  • huggingface/speech-to-speech — Modular VAD→STT→LLM→TTS voice agents with an OpenAI Realtime-compatible WebSocket.

  • microsoft/VibeVoice — Open-source frontier voice AI: long-form ASR, realtime TTS, BitNet edge engine.

  • 1jehuang/jcode — Hyper RAM-efficient coding harness for multi-session agent workflows.


▶️ YouTube

  • Kimi K3 Just Broke The Economics Of AI — Two Minute Papers | Short, paper-flavored take on how open-weight Kimi K3 punches cost/performance assumptions.

  • I’m disappointed — Matthew Berman | Nvidia’s letter, what “open source” means for models, open weights, and Anthropic’s stance in one thread.

  • Which business will “win” AI? — Matthew Berman | Qualcomm’s bet on infrastructure (data centers, autonomy, robots) rather than the model itself.


💬 Community

COMMENTS