📰 News
-
Astra flew a surveillance drone and ran a vending business solo — Andon Labs. On a one-year vending sim it averaged $15,515; Claude Fable 5.1 averaged $5,422. Astra’s worst run still beat Fable’s best. On Drone-Bench it became the first model to beat the human-AI baseline on all five tasks, from 3D reconstruction to person tracking. End-to-end success in one go: 2.8%. → The Decoder
-
Altman, Musk, and Hassabis backed Amodei’s in-lab auditors — Put independent evaluators inside the frontier labs. Google’s Peyman Milanfar pushed back: recursive self-improvement loops are unstable by nature, so a speed limit is unnecessary. “Stability is the speed limit.” → The Decoder
-
Iris-mini and Iris-pro lead open-weight search agents — AllSpark. Qwen-based 35B and 397B, 256K context. Training questions are reverse-engineered from the web’s link graph, so the agent has to chain steps instead of grepping. Tops its class on BrowseComp. Search training also transferred to tool use and office work. → The Decoder
-
Banning AI in class left students worse off — Two-year study at VU Amsterdam Law. The no-AI group finished last both years. Prof. Schrepel: “I was wrong.” Unguided ChatGPT still beat a ban. Berkeley Law is still banning AI from almost all graded work. → The Decoder
-
ElevenLabs shipped Music v2.5 on app and API — In 47,885 blind pairs, listeners preferred it most on R&B, hip-hop, and orchestral. Free tier: five lossless tracks a day; Pro: 400 a month. The UMG deal is for a later product. Trained on licensed stems, unlike Suno. → The Decoder
-
IBM’s Granite time-series model tops zero-shot under a commercial license — PatchTST-FM-r2, 385M. Best permissive-license model on GIFT-Eval. Apache 2.0. → Hugging Face
🔥 GitHub Trending
-
asgeirtj/system_prompts_leaks — System prompts from Claude Fable 5.1, GPT-6 Astra, Gemini, Grok. The hidden rules before the first reply.
-
jihe520/MathModelAgent — An agent for math modeling contests. Solves the problem and writes a submission-ready paper.
-
alsk1992/CloddsBot — A Claude-powered trading agent. Scans 1,000+ prediction markets and crypto venues, then executes.
▶️ YouTube
-
Deepseek did it again… — Matthew Berman. DeepSeek V4.1 Flash. Speed, again.
-
DeepSeek Is INSANELY Fast — Matthew Berman. Same Flash, as a short you can feel.
-
DeepSeek Fails the Rubik’s Cube Test — Matthew Berman. Fast is not the same as solving a cube.
-
The $1 Million Fluid Problem — Matthew Berman. A million-dollar prize on a fluid-dynamics puzzle.
💬 Community
-
Generating running routes with GPT-6 Astra — Simon. 5K and 10K loops from OSM in 27 minutes. The code vanished after compaction.
-
We Must Pace the Frontier — Amodei’s post. The speed-limit pitch itself.
-
Better AI code comment detector — Logistic regression on LLM-written comments.
-
ML on a Guitar Hero controller — A model that strums. Hobby turned experiment.