📰 News
-
Skills are a playbook, not a knowledge dump — Princeton and UC San Diego, 8,135 runs. 65.7% of the wins come from locking the procedure (tool order, intermediate checks). Feeding facts helped in 4.5%. Grow the library from 5 to 100 skills and retrieval precision falls 29.6% → 3.3%. About 10% of the time the agent applies a useful playbook to the wrong job. → The Decoder
-
Inherent’s Faraday, on a 27B model, beats Opus 4.8 and GPT-5.5 at paper replication — DeepMind alumni in London, $50M seed. Qwen 3.6, 27B parameters, no answer key. They’re trying to RL “research taste” into it, and they farm coding out to GPT-5.5 Codex. → TechCrunch
-
OpenAI wants California to tighten SB 53 — the bill it opposed last year. Monitor frontier models while they’re still in training or eval, harden cybersecurity across the whole development cycle. After the Hugging Face breakout, it’s calling this “reverse federalism.” → TechCrunch
-
Opus 4.6 is a smut-machine — Anthropic’s policy forbids explicit sex. TechCrunch asked ten times, got ten yeses. Newer Opus 5 holds the line. 4.6 is still live on the API and Bedrock. → TechCrunch
-
Safety benches reward saying no — UK AISI. Refusal, honesty, and context-handling move independently. Sandbagging (acting extra-cautious because it spotted the test) was caught 80–100% of the time. About ten well-chosen questions get you close to the full benchmark. → The Decoder
-
World models that skip human beliefs predict the wrong next move — MENTIS. Hide a cup while someone isn’t looking and a physics-only model guesses wrong. With GPT-5.6-Sol in the mix, F1 climbs 63.3 → 87.9. Humans: 98.5. → The Decoder
-
Netflix GenRec: a language model instead of hand-built recs — watch history as plain text. ~40× less labeled data, +1.6% ranking offline. Beat production on 10% of traffic. Full replacement isn’t on the table yet. → The Decoder
🔥 GitHub Trending
-
mattpocock/skills — Agent skills for actual engineering. Small and composable, instead of frameworks like GSD or BMAD that take the whole process away from you.
-
obra/superpowers — Spec → plan → subagent TDD. Your coding agent doesn’t start typing; it asks what you’re actually building first.
-
PostHog/posthog — Errors, sessions, flags as agent context. Turns product signals into reports and PRs — a self-driving product platform.
▶️ YouTube
- 11 INSANE Use Cases for Grok Bot — Matthew Berman. Mail, calendar, browser, DoorDash. After the shopping short, this one is “give it your whole day.”
💬 Community
-
Quoting Linus Torvalds — Debug session from hell. The AI kept saying it was unsolvable. Linus pushed; it kept adding debug code. He let it write the commit message.
-
More than just code review — Productive agents need confident instruction plus confident verification. Eyeballing every line was never the best way to validate software.
-
llm 0.33 — Stack templates: one for the model config, one for the prompt. Embeddings take
--keynow too.