Events

🇬🇧 HoloRun in London (October 16) – H Company will host a 3.5-mile run at an easy, chatty pace at 7:00 AM, then breakfast and nerdy talk.

🇫🇷 Beyond agents in the life sciences in Paris (December 2) – Orakl Oncology, Opsin and TychoBio will show how to validate AI agents in biology, over breakfast during Bioweek 2026.

Audio

🗣️ ElevenLabs released Eleven v4 voice models – A new flagship text-to-speech model and a faster Turbo variant, billed as the company’s most emotive and quickest voices yet. Both arrived with a showcase demo video and took the top spot on the Artificial Analysis ranking.

⚡ Gradium cut text-to-speech latency below 50ms – Voice agents get roughly 200 to 300ms to reply before a pause feels unnatural. Gradium’s new default model starts speaking in about 50ms, faster than ElevenLabs v4 Turbo, without trading away naturalness or its handling of phone numbers and codes.

📝 Microsoft released MAI-Transcribe-2 speech recognition – Microsoft’s second speech-to-text model tells speakers apart, timestamps every word and can be steered toward domain jargon and names across 60 languages. It posted the lowest word error rate on FLEURS and turned an hour of audio into text in about 10 seconds.

🪶 Phonon-2 shrank speech recognition to 164MB – An open-weight speech-to-text model from Fermion that averaged better accuracy than OpenAI’s Whisper large, a model ten times its size. It transcribed an hour of audio in 20 seconds on a MacBook Air and shipped under CC-BY-4.0.

Autonomous Agents

🕵️ Meta Muse ignored Mac Messages permissions – Meta’s Muse agent, built to research, brief and shop from a dedicated virtual computer, synced 187,000 lines of a journalist’s Apple Messages history to the cloud without permission, even with Full Disk Access switched off.

☎️ Meta Muse calls came from humans – Meta promoted Muse calling businesses to book reservations, but internally tested handing those calls to trained human agents in call centers. Testers learned only afterwards that a person had called, and staff warned training is not a security mechanism.

🥧 Earendil shipped Pi 1.0 agent harness – While agent tooling churns weekly, this minimal, extensible harness adopts features only once they prove themselves. Version 1.0 added native MCP support, deferred tool loading and Pi Durable for long-running agents, leaving sandboxing to Docker or extensions.

Biotech, Health, and Chemistry

🧬 DeepMind’s SynthID Bio watermarked AI proteins – Arguing that tracing AI-designed proteins matters for biosecurity, the Nature paper hid detectable watermarks in AI-generated protein sequences and AlphaFold 3 structures. Watermarked binders kept their binding strength in lab tests. https://github.com/google-deepmind/synthidbio

  • Community take w/ Félix Raimundo of TychoBio: “It solves a problem no one cares about. Google even managed to get some of the open peer review censored, and the reviewers’ questions are especially not nice.”

🧪 Joe Horsman put atoms before bits – AI got good at generating biological hypotheses, but the hard part of science is validating them. The essay argued AI biotechs still have to build labs, generate proprietary data and run trials, as whoever does the hard work captures most of the value.

Image, Video & 3D

✏️ Ideogram 4.5 fixed multi-turn image editing – With each round of edits, leading image models pile up artifacts, pixel shifts and color drift. Ideogram 4.5 removed that buildup, making long editing sessions usable. Live in Ideogram, the API and partner apps, with open weights promised soon.

🎭 Tavus Griffin claimed video Turing test – Griffin holds live, two-way AI video conversations, and Tavus claimed 48% of people who talked to it believed it was a real human, versus under 3% for earlier systems. That figure allegedly came from its own study of 54 one-minute calls, not an independent protocol.

🧱 LEGO-Anything rebuilt scenes as Blender code – Rather than output a fixed 3D mesh, coding agents write and run Blender code to rebuild a scene from one photo, compare renders with the image and revise, leaving an editable scene. The best agent, GPT-6 Astra, scored 53.4% indoors and 39.6% outdoors.

  • Agents failed in three recurring ways: weak opening scenes, edits that broke what already worked and unreliable judgment of their own progress.

  • A training-free plugin targeting those three issues, including version control to protect correct progress, lifted all six tested models by up to 62.7%.

  • Detections, masks and depth read off the rebuilt scenes worked, but fell well short of specialized vision models.

🖼️ FLUX 3 Image brought pixel-level control – Instead of hoping a prompt lands the layout, users place each element in a box on a canvas, then edit a finished image box by box while everything untouched stays put. An LLM can plan the layout itself, which made the model a fit for agents.

🤠 Refresh’s BlenderBench tested agents on animation – Going past 3D modeling, this benchmark hands agents an unrigged cowboy mesh and asks them to build its skeleton, bind it and animate a tightrope balance from a reference. Claude Opus 5.5 was capped at 0.40 for misplaced joints and feet that never stepped.

Cyber

🔐 OpenAI insider looked past the sandbox – Writing in a personal capacity, a member of OpenAI’s Agent Security team argued this year’s incidents followed capability jumps that shocked insiders, from Millennium Prize math to cyber and swarming, outrunning how fast security culture can mature.

Language Models

🎼 Anthropic released Claude Sonnet 5.5 – Anthropic’s workhorse model for well-scoped everyday jobs, like fixing bugs or building polished documents, slides and spreadsheets. At the same price as Sonnet 5, it worked noticeably faster and cheaper, because it got the same job done with far fewer tokens.

  • On many tests it came close to Opus 5.5, but Anthropic said its flagship stays clearly better at open-ended work that needs sustained judgment.

  • Early testers saw it get there in fewer steps: Base44 needed half as many attempts per app build, and Balyasny’s finance answers used about a quarter of the tokens.

  • Its cybersecurity skills jumped to the level of the previous Opus, making it the first Sonnet shipped with Opus-style safeguards, with risky security requests visibly handed back to Sonnet 5.

🛡️ Google DeepMind introduced Gemini 4 Argon – Google’s new frontier model is built to stay on hard, long jobs, from big software projects to legal and finance work, and can produce far longer answers in one go than before. It was also trained to hunt down and fix security holes on its own. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/

  • Inside Google, teams of Argon agents rewrote old C and C++ code into safer Rust, including a video decoder that came out memory-safe and nearly three times faster.

  • Security firm Wiz used it to uncover a critical flaw exposing sensitive personal data in hospital software worldwide, one earlier frontier models had missed.

  • Google is still adding safety checks, including monitors that read the model’s reasoning and can stop it mid-task, before opening it to developers and paying subscribers.

☀️ OpenAI released GPT-6.1 Sol – OpenAI’s mid-tier model now handles most of the demanding coding, computer use and office work its flagship GPT-6 Astra does, for about a fifth of the price. It matched Astra on real software engineering tasks and read complex PDFs better than Opus 5.5.

  • The biggest jumps came outside coding: it got noticeably better at operating a computer and doubled its score on scientific tasks.

  • Reused context now costs 95% less than fresh input, so agents can run longer on the same budget.

  • For users who want speed over savings, OpenAI added up to 8x faster versions of Astra and Sol, reserved for a new $500-a-month Pro tier.

🎯 InternLM’s Intern-Decision-2B took on Jev – A Jev-style decision model fine-tuned from Qwen3.5-2B: given a shared state, typed questions and optional images, it scores every allowed answer in a single pass instead of generating text. Training code, calibration and a browser demo came with it.

  • Community take: Gabriel Olympie of 2501.ai called it “the first open source Jev variant that actually beats Jev. Took them less than a week to reproduce this marketing breakthrough”, while Clement Poiret of Rhizome Labs added: “Jev replicated by OpenAI based on Luna, typesafeai broke the record of time-to-deprecated.”

🔤 Fastino open-sourced GLiNER2.5-Decide encoder – An open alternative to Jev that skips the LLM: a 340M-parameter encoder answers typed questions about a text with probabilities and enforces rules linking the answers. It runs on a CPU in 167ms, suits air-gapped setups and beat a Qwen3.5-4B decision model.

🔀 Cloudflare open-sourced Clef decision models – Agents make constant small calls, like routing a ticket or flagging a phishing site. Clef and Clef-flash answer them with typed outputs instead of open-ended text, as a Jev API-compatible pair that adds image input and a 64k context, hosted on Workers AI.

  • Clef runs on Qwen3.8-27B and Clef-flash on Qwen3.5-9B, with a final step that reads the answer straight from the model’s internal state rather than generating text.

  • Cloudflare’s threat intelligence team used it to fetch, render and classify a website in 2.2 seconds, against 4.7 seconds for its fastest general LLM.

  • A reinforcement learning fine-tuning platform launched alongside, offered through Cloudflare’s forward-deployed engineers before a planned self-serve version.

decision-index-vs-latency.png

📂 Context Language Models rewrote their context – Instead of a hand-built harness deciding what stays in the context window, the model treats its context as a file it can append to or freely rewrite with Bash. It beat engineered context-management strategies on accuracy while using less compute.

  • On BrowseComp-Plus deep research it scored 11.4% higher with 21.5% fewer FLOPs, and on 12-hour EdgeBench it gained 5% with 59% fewer FLOPs.

  • A Suffix Cache Reuse mechanism kept cached computation usable after context edits, cutting server compute by 35% versus standard SGLang.

  • Models invented their own tactics, such as trackers for multi-agent orchestration, new chat roles for internal notes and reusable context-management functions.

🐜 Ant Ling announced Ling-3.1-flash – Ant Ling’s new model targeted work, coding and healthcare tasks with a context window of up to 1M tokens, activating only about 25B of its roughly 560B parameters per token. It scored 75.16 on FrontierSWE, with open weights promised soon.

Programming

🎨 How do you work with designers – With AI blurring design and code, how far should designers go: hand over mockups, ship front-end code on dummy data for engineers to wire up, or vibe code the backend tweaks too? Most teams landed in the middle, and project size decided the rest.

  • Charly Poly of Browserbase, a 70-person SF team where brand is central to the marketing strategy: “We are heavily using Claude Code and Claude Design to create assets across all teams: sales decks, website graphics, full websites, product parts. Designers create skills or design systems for static and animated assets that are given to the teams, so I’d say 80% of design work is self-serve here now. Interestingly, we are hiring more designers right now, as the remaining 20% still requires more and more work, and a human in the loop for both messaging and overall brand consistency. On designers shipping code, we have a dedicated Design Engineer who handles both design and code.”

  • Jérémie Bordier of XHR, which serves millions of users and keeps a senior UX designer and a senior PM, both much faster with AI: “Our engineers feel they need less guidance from product and designers, but so far we’re pushing back, because keeping the PM and designer in the loop makes sure we actually answer user needs better and avoid half-good solutions we have to rework after the feature ships. Going from the mockup to the code is pretty straightforward with all the knowledge and design system Claude is aware of. At the end of the day, each keeps their expertise, but we need way fewer junior people across the board.”

  • Nikolay Tchakarov of Asteria, with 2 designers and 5 engineers: “The distinction is definitely getting blurred, so I’d say we are level 2.5, but it depends on the kind of features. Small frontend adjustments get implemented directly in the real codebase. Designers use Claude assets to iterate quickly, and passing those as context to re-implement within the actual codebase works well for small features. For larger projects we are seeing the limit of it though: code architecture choices become more important, and that’s where the limit of design-driven features is most visible. We’re now wondering how to streamline the handover for large projects in the AI agents era.”

  • Lior Oren, on a project with 6 engineers and 2 product people: “The ideal setup is designers owning a design system that’s translated 1:1 into a component library. These are the building blocks for a unified UI, simpler development, and much easier work with AI coding harnesses, though Codex will try to breach these rules like a developer on the last day of a sprint. With a good component library, designers can already ship static versions of features. With a perfect one, they can go to level 3, minus the backend part. We used to also have an interaction designer, but those profiles are really hard to find. That’s probably our biggest UX gap today: product flows.”

  • Robert Hommes of Moyai.ai: “We went fully to product engineers, so usually one person does the mockup or first implementation, depending on what works best, and then ships it.”

  • Denis Brulé of Finegrain: “ChatGPT is really good at it these days. It produces great-looking designs. A simple PRD wrapping those can then be fed to Claude to automate the dev work. A huge productivity boost.”

🛤️ DHH said 37signals stopped handwriting code – Calling Opus 4.5 the Kodak Brownie of software, DHH said hand-written code at 37signals is now treated like a bug: when an agent fails, the team fixes the factory. He wrote 150,000 lines of code in August, about 60 times his long-run average.

  • Designers vibe coding Basecamp 5’s final features in the spring left the architecture “like a Swiss cheese”, but DHH said going back to manual review was the wrong call: had they waited for Fable, it would probably have worked as intended.

  • HEY is being rebuilt as six native apps, with its backend rewritten in Rust by agents for 99% less CPU and 95% less memory, light enough that peak traffic could likely run on a Raspberry Pi.

  • His one prescription: every app should ship a CLI “by next Friday” so users can bring their own agents, instead of bolting a chatbot into each product.

Reinforcement Learning

🎓 Agentic-DPO and Thinking Machines’ on-policy distillation – Gabriel Olympie of 2501.ai picked these two papers for an RL pipeline that transfers agent skills from a large teacher model to a smaller student, after plain fine-tuning (SFT) and standard DPO proved inefficient for the job.

  • Agentic-DPO: at each step of an expert trajectory, the student proposes its next action and DPO teaches it to prefer the teacher’s. On τ-bench it lifted a 9B model from 21.7% to 41.4%, matching online GRPO without environment interaction.

  • On-Policy Distillation: the student runs the full agent behavior, the teacher grades every token in a single forward pass, and training minimizes the KL divergence between the two.

Robotic, World AI

🌍 InSpatio-World 1.5 turned photos into worlds – Feed it one image, several images, a panorama or a video, and this open-source model builds a scene to roam in real time, revealing areas the original camera never saw and, with video, moving freely through time as the action unfolds.

  • Community take w/ Julien Millet of United Bits: “Everybody says they are making a world model when they’re just doing spatial coherence or video extrapolation.”

Contributors This Week

Gabriel Olympie, Gabriel Duciel, Félix Raimundo, Jérémie Bordier, Charly Poly, Julien Millet, Kemal Toprak Uçar, Louis Choquel, Quentin Dubois, Amine Saboni, Clement Poiret, Maziyar Panahi, noeachache.1, Robert Hommes, Denis Brulé, Jules Belveze, Julien Seveno-Piltant, Lior Oren, Nikolay Tchakarov, Pierre Chapuis