📰 News
-
GPT-5.6 Sol disproves a 30-year stats conjecture in ~90 minutes — A UPenn researcher used Sol Pro to show Benjamini–Hochberg FDR can miss its target on correlated normal data. GPT-5.5 failed after 20+ hours. → The Decoder
-
OpenAI’s GPT-Red beats human red teamers — Self-play attacker finds prompt-injection-style flaws in 84% of tests vs 13% for humans. Sol is 6× more resistant to direct injections than four months ago; ~3.8% of stronger attacks still land. → The Decoder
-
Codex encrypts agent-to-agent instructions — Inter-agent delegation is now opaque in logs. Mandatory on Sol/Terra. Privacy move or rival-training block? Some flaky handoffs reported. → The Decoder
-
$230 Codex Micro keyboard amid Apple hardware fight — Limited Work Louder collab: Agent Keys, reasoning dial—a desk command center for agent fleets. Screenless ChatGPT device still in the rumor mill; Apple’s suit continues. → TechCrunch
-
Bonsai 27B: open reasoning that fits on an iPhone — PrismML compresses Qwen3.6-27B to ~3.9 GB (1–2 bit) with ~90% quality. Bet: local agent loops, zero cloud token cost. Apache 2.0. → The Decoder
-
Anthropic + Blackstone’s Ode: implementation, not just models — $1.5B JV, ~100 forward-deployed engineers, Claude-first, pilot-to-production gap. Same category as OpenAI’s Deployment Company. → TechCrunch
-
Inkling + Meta layoff suit — Thinking Machines’ ~975B MoE multimodal (1M context) hits HF. Meta workers claim AI systems ranked people for May’s 8k cut. → HF · The Decoder
🔥 GitHub Trending
-
Shubhamsaboo/awesome-llm-apps — 100+ runnable agent, RAG, and skill templates.
-
mattpocock/skills — Small composable skills for real engineering work. skills.sh + Claude Code plugin.
-
Dicklesworthstone/destructive_command_guard — Blocks destructive git/shell commands before agents run them.
-
virattt/ai-hedge-fund — Educational multi-agent “hedge fund” with investor personas. Not for real money.
-
Nutlope/hallmark — Anti-AI-slop design skill: audit/redesign verbs.
▶️ YouTube
-
Claude’s Brain Has A Secret… — Two Minute Papers | Quick tour of Claude internals research.
-
OpenAI vs Anthropic — Matthew Berman | Pricing, usage, positioning.
-
Deepseek ban? · China’s chips — Berman shorts on regulation and domestic silicon.
💬 Community
-
How I tricked Claude into leaking secrets — Nested links bypassed web_fetch’s user/search-only URL guard. Name, city, employer extracted; hole closed.
-
AI Data Centers and the Concentration of Wealth — Schneier on who captures the surplus when compute capital piles up.
-
MiMo-V2.5 inference pipeline — Xiaomi’s end-to-end serving notes.
-
Inventing ELIZA — Open-access history of the first chatbot—useful mirror for agent hype.