📰 News
-
Moonshot drops Kimi K3 weights and infra — After stirring the frontier race, China’s Moonshot put Kimi K3 on Hugging Face and open-sourced pieces of the stack: attention kernels, an MoE comms library, agent-scale tooling. Claim: ~2.5× intelligence per compute. Scores sit near Fable 5 and GPT-5.6 Sol, but cyber and math lag — distillation rumors included. → The Decoder
-
Claude shared chats and Artifacts showed up on Google —
site:claude.ai/sharesurfaced public share links. Reports of medical records, internal docs, even kids’ contact info. Anthropic says links only get indexed if posted somewhere crawlable; by Monday afternoon Google results were gone. Same class of mistake as last year. → TechCrunch · The Decoder -
Microsoft ships MAI-Cyber-1-Flash — Paired with the MDASH multi-agent system, it hits 96% on CyberGym. Flash handles ~90% of work; hard cases go to GPT-5.4, with a claimed ~50% cost cut. Also: Perception, a real-time threat agent. Preview Nov 3. → TechCrunch · The Decoder
-
METR’s “expenditure horizon” for agents — Dollar point where AI and humans cost the same for the same gain, measured on the NanoGPT speedrun. Humans ≈ $2,500 per 1% speedup. Older models barely move the needle (horizons $0–$3.3k). Opus 5-class models not in the paper yet. → The Decoder
-
OpenAI: more workers use ChatGPT for other people’s jobs — “Task crossover”: of 800k+ work messages, 43.5% of job-specific queries were for another profession. Marketing and engineering cross most. Stronger at small companies. → The Decoder
-
Delhi High Court rejects ANI injunction against OpenAI — Evidence articles post-dated training cutoffs; judge leans RAG over memorization, no verbatim copies. Tentative fair-use view on training + public benefit. Main case continues. → The Decoder
-
HF breach reignites alignment vs. containment — Sandbox bug or a model that wanted to cheat? GPT-5.6 Sol looks worse than 5.5 on agentic misalignment metrics. OpenAI patches and monitors; it isn’t slowing capability work. → TechCrunch
🔥 GitHub Trending
-
block/buzz — Self-hosted workspace where humans and agents share one Nostr-shaped event log.
-
citrolabs/ego-lite — Browser that splits your tabs from agent Spaces while sharing logins (Codex, Claude Code).
-
pingdotgg/t3code — Minimal web GUI for Codex, Claude, Cursor, OpenCode. Try
npx t3@latest. -
OtterMind/Chat2DB — AI SQL client for 30+ databases; NL→SQL, MCP CLI.
-
CoreBunch/Instatic — Agentic self-hosted visual CMS that ships clean static HTML.
▶️ YouTube
- AI can’t READ this — Matthew Berman | Short on visual text / CAPTCHA-style limits models still hit.
💬 Community
-
Open Weights and American AI Leadership — Microsoft’s open-weight stance, sharper now that it’s shipping its own security models.
-
A tour of MLIR — Dialect stack tour under modern ML compilers.
-
Languages as designed latent spaces — Programming languages as designed latent spaces — a neat framing for the LLM era.
-
Inside the token reseller relay market — Simon on discounted-token proxies and key abuse. Hard API spend caps, please.