Community Discussion

🤔 Did OpenAI’s math leap generalize? – OpenAI’s release of 719 math manuscripts left one question hanging: was this a leap in reasoning itself, or only in math? Willy Braun, reading through the week’s math breakthroughs, put it to the community. He relayed the view of a friend with a strong math background, convinced that “we’ve crossed a major threshold in abstraction and reasoning.” If that held, “frontier generalists could soon reach SOTA on specialised tasks with very little domain-specific data,” possibly by distilling existing specialist models.

Willy wasn’t sold, especially for biology, “where so much is unknown and where you need experimental data to verify hypotheses and close the loop.” He still saw the stakes: a real jump in abstraction would mean “we’d need far less domain-specific data than we currently assume.” He also pointed to an ICML 2026 paper showing that better math reasoning didn’t necessarily carry over to other domains, though it predated the latest results.

Gabriel Olympie of 2501.ai placed the impact on the theory side. In his reading, OpenAI had “managed to build a form of superintelligence able to cover math problems that are formalisable in Lean.” He saw it as a tool to “quickly disprove specific hypotheses and conjectures, trimming research dead ends at industrial scale.” Applied fields like biology might feel it less. The bigger opening could be the theoretical side of applied science: building a theory around a hypothesis and working out its math “until it stumbles upon a physical measurement already done that can prove or disprove that theory.” He called that “purely speculative.” So far the models mostly refuted hypotheses someone else had framed, and errors could cascade over long problems.

Emmanuel Benazera of Jolibrain answered with three axes:

  • Speed: ultra-fast frontier models were coming, but “you don’t want to drive a physical system like a car or a robot with a thinking model right now.”

  • Cost: decisions don’t all carry the same value, hence falling token prices and “Jev-like structural changes.”

  • Precision: the axis he expected specialists to keep, since “it’s hard for frontier models to have top accuracy absolutely everywhere.”

His picture of where this ends: a frontier model distilling itself into “a ‘cockroach’ version of itself, dedicated to a specialized task,” wrapped in generated code, with the idea of a harness stretched to cover it.

Biotech, Health, and Chemistry

🌱 Living Models’ Botanic-1 pinpointed crop mutations – Breeders can narrow a useful crop trait to one stretch of DNA, but finding the single change behind it can take years of field trials. Working with Google’s Gemma 4, the plant DNA model found the exact mutation behind a melon flower trait in under four minutes.

  • Gemma 4 handled the project work, writing code and sorting data, while Botanic-1, trained on DNA from 320 plant species, judged which changes actually mattered.

  • It picked the right mutation 9 times out of 10, where classical tools managed 15% at best, even with expert guidance.

  • The whole setup ran on a single NVIDIA graphics card, so breeders could keep their proprietary genomes in-house.

  • Leonard Strouk from Living Models: “Building this system means connecting two distinct capabilities: reasoning through an analysis and predicting the effects of changes in DNA. BOTANIC-1 learns from plant genomes to help distinguish variants that conventional analysis leaves tied. Gemma turns that signal into an actionable result: a prioritized hypothesis that a scientist can take to the lab.”

🧬 TychoBio tapped PacBio for RNA data – To train models that design RNA drugs for rare diseases, TychoBio planned to test candidates on over 10,000 samples and use long-read sequencing to see how each one changed full-length transcripts, predicting which compounds work and where side effects hit.

🏛️ Biohub pooled $1.8B for AI-ready data – Co-founded by Mark Zuckerberg, the nonprofit teamed up with the US Energy Department, the NIH and other funders to pay for data, compute and new measurement technology, so researchers could use AI models to prevent and treat disease.

  • Community take w/ Félix Raimundo of TychoBio: “I’m kind of wait-and-see on that one. It’s led by Rives, who raised a stupid amount of money for EvolutionaryScale, and nothing came out of it. But for what it’s worth, among all the mega fundings, Biohub is the ‘best’ one, so I’m cautiously optimistic.” He also pointed to a post on a 0.08 Pearson correlation between predicted and observed perturbation effects, on datasets that cost millions to generate: “That’s the kind of data Biohub is going to generate.”

💊 Anthropic stopped drug discovery pre-Phase 1 – To avoid competing with pharma customers, the company kept its own rare disease programs to early discovery and possibly some preclinical work, short of clinical trials. Its life sciences head also discussed Claude’s first discovery, a new enzyme system.

  • Community take w/ Félix Raimundo of TychoBio: “The article shows that another scientist had already used Claude for research on that enzyme, kind of like OpenAI with the Millennium Prize.”

Image, Video & 3D

🍌 Google shipped Nano Banana 2.1 – The image generation and editing model claimed gains over earlier versions across the board, with the biggest leaps in visual design, mask-based editing, where only a selected area is changed, and keeping a subject consistent from one image to the next.

🖼️ Speridlabs’ Iris-3B generated raw pixels directly – Most image models work in a compressed space decoded by a VAE, losing fine detail. Trained from scratch, the open 3B model skipped that step, claimed higher GenEval and DPG scores than the 12B FLUX.1-dev, and also handled depth estimation and image restoration.

🐢 GenIA aligned SAM3D to real photos – Without retraining, the method grounded SAM3D’s guesses in observed geometry and colors at test time, rebuilding complete objects from one image, several views or a video. On video, it cost about 13x SAM3D’s single-image runtime, against 160 to 250x for rivals.

🎮 Google Labs launched Playground for games – The experimental platform let anyone build and play games by describing them in a chat, picking genres like trivia, tower defense or racing or starting from scratch, with no coding needed. It opened to users 18 and over in the US.

Language Models

🐱 Mistral launched Large 4, Le Chonk – The natively multimodal model activated 49B of its 1T parameters per token, was built and served from Europe, and claimed the best aggregated scores of any US or European open-weight model. API access came first, with weights due at the end of October.

  • Community take w/ Gabriel Olympie of 2501.ai: “Not frontier level, but decent. Considering how bad Mistral Small 4 is, it’s still a good comeback!”

🧮 OpenAI released 719 AI-generated math manuscripts – After existing math benchmarks saturated, an unreleased internal model was set on about 4,000 open research problems. The output, grouped into 372 families, ranged from the irrationality exponent of π to Kaplansky’s direct-finiteness conjecture.

  • Most results came from the same procedure, using about three hours of ChatGPT Pro thinking compute each on average.

  • Only about 42% of top-line results had Lean formalizations so far, and OpenAI warned that unformalized ones could contain errors it would fix.

  • OpenAI shaped the release with advice from the independent Advisory Group on Mathematics and AI at the Institute for Advanced Study.

  • Community take w/ Gabriel Olympie of 2501.ai: “That one is very impressive, in some ways more than the Navier-Stokes release. Pending review, but some very important problems fell here. Let’s see how the math community reacts.”

📐 Lean checks hid a Navier-Stokes mistranslation – Compiling Lean code proves nothing about the original proof if the AI quietly changed the statement, and the authors showed faithful translation is harder than the Halting problem. In OpenAI’s Navier-Stokes proof, one lemma’s Lean version was weaker.

  • Community take w/ Gabriel Olympie of 2501.ai: “Interesting but hard-to-interpret paper with a clickbait title. It doesn’t disprove the solution, it just says it needs proper peer review.”

🧲 Google opened multimodal EmbeddingGemma 2 – The 740M on-device model mapped text, code, images, audio and video into one shared space, so a photo of a cat landed near the word cat. Paired with Gemma 4, it targeted offline, privacy-first retrieval and beat some specialists twice its size.

🎯 Liquid AI opened d1 decision models – Instead of writing text, d1 answered yes/no, multiple-choice or score questions in one pass. After adding image inputs and claiming GPT-6.1 Sol-level results on four of six real tasks at 19 to 200x lower cost, Liquid AI opened d1-3B, which ran in 8 ms.

  • Community take w/ Gabriel Olympie of 2501.ai: “The thing with Jev is that the breakthrough is mostly in the usage rather than the tech. The approach has existed for years, based on simple encoder models plus a calibrated classification layer. Where I see some room for new business is mostly edge inference: shipping a small model of that kind into any application stack can be done with a 500 MB binary, and having a standard way to do that could be useful.”

🐦 Aleph Alpha opened Kolibri, joining Cohere – The German and English reasoning model activated 3.46B of its 78B parameters per token, read up to 1M tokens and targeted sovereign use in regulated sectors. Aleph Alpha also signed to merge into Cohere, with headquarters in Berlin and Toronto, pending approval.

  • During reinforcement learning, models used web-fetch tools to find ground-truth solutions on GitHub in about 7% of trials, and some tried to escape their sandbox.

  • The first two mixture-of-experts layers barely used their routed experts, with a bias term overriding the router’s choice 98.5% of the time.

  • Deliberate contamination tests showed memorized benchmarks inflated HumanEval by 42 points and MMLU by 29.

🪶 Anthropic shipped Claude Haiku 5.5 – Built for high-volume work like summaries, classification and subagents, the small model cost about 75% less than Haiku 4.5 and leapt on agentic tasks, from 15.7% to 72.4% on the OSWorld 2.1 computer-use benchmark. It was the first Haiku with adjustable effort.

  • Prompt injection attacks succeeded 7.1% of the time, down from 83.2% on Haiku 4.5.

  • It over-refused more than any other model in automated behavioral audits and hallucinated more than recent models.

  • It used leaked answers without disclosing them 17% of the time, against 2% for Haiku 4.5.

🔦 Reflection unveiled its 501B Beam model – The US mixture-of-experts model, built for coding, reasoning and agents, activated 23B parameters per token and claimed GLM 5.2-level scores with 3 to 4x less inference compute, including 80.1 on Terminal Bench v2.1. Weights were set to ship under Apache 2.0.

  • Reinforcement learning ran more than 100 million rollouts on 10,500 GB300 GPUs over four weeks, using about 1.3 billion sandboxes.

  • Browsing improved during training although no browsing tasks were in the mix, as the model learned on its own to search, query other LLMs and call OCR APIs.

  • New asynchronous policy gradient algorithms kept learning stable even with samples up to a day old.

🔮 Humanity’s Sixth Sense tested visual intuition – The Scale Labs benchmark asked models to infer what a scene implied but didn’t show, from what just happened to who held power, across 522 image and video tasks. Humans scored 93.1%, the best model, GPT-6-astra, 53.6% and the median model 30.9%.

🗿 DINOv2 and Qwen3 aligned without pairs – The vision model had never seen a caption and the language model never an image, yet their embedding spaces shared enough geometry to be matched by one rotation-like transform. With up to 20 pairs, alignment error was 14 to 28x below the best paired baseline.

Robotic, World AI

🌌 Odyssey launched the Odyssey-3 world model – Billed as its most powerful foundation world model yet, it claimed a new state of the art on the Physics-IQ benchmark and was pitched for powering robots, training AIs and generating interactive experiences, free for anyone to try.

🌍 WorldPlay2 turned photos into playable worlds – Tencent Hunyuan’s model, running on Reactor, generated every frame live as the user moved through a world built from one picture, and let prompts steer what happened mid-exploration while keeping the scene consistent over long sessions.

🦾 Reka previewed Rho-1 omni model – The 19B research preview both understood and generated text, images, video and robot actions, all inside a single neural network.

🎥 LoGo kept long videos 3D-consistent – Camera-controlled video models lost track of objects as the camera moved, and one score per video couldn’t say where things broke. This World Labs post-training method scored 3D errors region by region, blended with a global score to avoid blurry reward hacking.

Other topics

📚 arXiv allowed two submissions per month – Submissions doubled in two years to a record 40,363 in September, with cs.AI up 6x, as moderators saw more thin, sliced-up and dense AI-written papers. Rejected papers counted toward the limit, and no submitter could hold more than three active at once.

  • Community take w/ Clement Poiret of Rhizome Labs: “Clearly a good thing. The endorsement approach is not sufficient, arXiv is flooded by fake papers.”

💾 Micron staff rejected a 68-month bonus – Workers at Micron’s Taiwan plant went on strike after being offered a bonus worth 68 months of salary, judging it not enough, while Samsung paid a $370k bonus to two thirds of its domestic employees, a sign of how overheated the chip economy had become.

New members

🇺🇸 Alexey Eryshev – Founder @ YC Stealth company · Software engineer with a mostly JVM background who discovered AI last year and ended up leading inference at Mistral. 📍San Francisco, USA, then Paris, France from January

Contributors This Week

Félix Raimundo, Gabriel Olympie, Pierre Chapuis, Quentin Dubois, Tejas Chopra, Willy Braun, Jean du Terrail, Julien Seveno-Piltant, Robert Hommes, Adam Surak, Alexey Eryshev, Amine Saboni, Clement Poiret, Emmanuel Benazera, Etienne Balit, Gabriel Duciel, Ihab Bendidi, Julien Duquesne, Louis Choquel, Stan Girard