Autonomous Agents

🦊 Shipfox launched agent CI workflows – An open-source, MIT-licensed platform defined agents as YAML in your repository, triggered by events like Sentry alerts or CI failures. Agents ran in isolated runners that mixed model steps with shell commands, retrying until your checks passed. πŸ‘‰ Access the Private Beta here!

Biotech, Health, and Chemistry

🧬 Merck and Moderna melanoma trial hit – A Phase 3 study of a personalized mRNA cancer vaccine combined with Keytruda met its main goals, cutting recurrence and distant spread versus Keytruda alone in resected stage 2B to 4 melanoma, the first such win for an individualized neoantigen therapy.

Image, Video & 3D

🎬 LTX 2.5 keyframe mechanics dissected – A hands-on breakdown by Jeremy Ringard from the community compared LTX 2.5 against MiniMax H3, with the findings largely confirmed by Mehdi Si-Mohammed and Julien Millet. The short version: the lighter distilled model is enough for most work, the real quality gains come from a multi-stage pipeline, and LTX 2.5’s most interesting trick barely exists in public tooling yet.

  • The heavier dev model barely beats the distilled one in normal use, so it is not worth routing through just for quality. The real gains come from multi-stage inference: a low-resolution pass, a latent upscale, then a full-resolution denoise. The dev model is still needed for training LoRAs.

  • On multi-shot clips, LTX holds a single camera shot until a change is actually needed, while H3 tends to cut and re-edit on its own once clips run past eight to ten seconds.

  • LTX compresses video harder than H3, packing frames in groups of eight versus four and shrinking each spatial dimension 32x versus 16x. That makes it faster and cheaper to run, but weaker on fast motion and fine detail.

  • LTX 2.5’s headline trick, called generated keyframe slots, has the model create its own keyframes in a first stage, store them uncompressed, and reuse them as references while denoising, with no user images involved. It adds roughly thirty percent more tokens on a 121-frame clip.

  • Despite being described in the docs, this internal mechanism does not appear to be exposed in any public tool. ComfyUI’s keyframe nodes only place user-provided images as conditioning, so almost nobody has actually implemented the real thing.

  • LTX still hallucinated more than expected for its reputation, with flipping fingers and objects passing through each other, which the group traced to the model having no real sense of 3D space. Julien worked around this by driving the camera with a Gaussian splat and feeding the rendered view in, so the model no longer guesses the geometry.

🧍 4DAnyone rebuilt humans in 4D – A method reconstructed photorealistic, free-viewpoint 4D humans from a single casual phone video by generating many consistent novel views and lifting them into Gaussian splats, using new tricks to prevent structural drift at scale.

  • Community take w/ Julien Millet: β€œThat is nice, but those are animated Gaussian splats. The amount of data must be pretty insane. Still, for some pipelines it could be fun.”

πŸ“· GNM Webcam Puppet ran in-browser – An open-source demo turned a webcam feed into a live 3D head fully in the browser, mapping 478 tracked face points onto Google’s parametric head model on the GPU, with nothing ever uploaded off the machine.

πŸ”„ LTX 2.3 LoRA warped camera views – An open LoRA for LTX-Video 2.3 regenerated a scene from a new angle given azimuth, elevation, and distance, using depth-warp conditioning and keyframed camera orbits for controllable novel-view shots inside ComfyUI.

Infrastructure

πŸ™ GitHub suffered a major outage – Network saturation on Central US load balancers took GitHub down for nearly eight hours, hitting Issues, pull requests, APIs, Actions, and Copilot. A misconfigured autoscaling policy and retry storms cascaded across regions.

  • Error rates peaked near twenty percent for web and API traffic and around fifty percent for archive downloads.

  • An Istio sidecar hit its concurrency limit and failed to scale, cascading until four HAProxy nodes exhausted their flow limits.

  • A retry bug in VS Code amplified traffic roughly tenfold, pushing one Copilot service from nine thousand to a hundred thousand requests per second.

Language Models

🐣 Qwen3.8-27B arrived as open model – Qwen released a 27B dense vision-language model with native image and video understanding, tunable thinking, and a 262K context extendable to a million tokens, posting strong coding, agentic, and research benchmark scores.

  • Community take w/ Gabriel Olympie (2501.ai): β€œIt is crazy powerful, though it overthinks slightly.”

🦀 Ornith-1.5 open family rivaled Opus – A new open-source LLM family arrived in 9B, 35B, and 397B sizes trained with self-improving methods. The 397B model claimed benchmark scores near Claude Opus 4.8 on reasoning, coding, and agentic tasks, and GGUF builds topped 100K downloads.

⚑ DiffusionGemma generated text in parallel – Google DeepMind fine-tuned a Gemma 4 mixture-of-experts model into a discrete diffusion language model that refines 256-token blocks in parallel, reaching about 1,500 tokens per second on one H100 while keeping thinking mode and multimodal input.

πŸ•΅οΈ Ox Alpha stealth model impressed – A stealth coding model on OpenCode and OpenRouter, rumored to be a new GLM 5.X from Z.ai, scored 80 percent on a ten-task DeepSWE subset, edging out GPT-5.6, Fable, GLM-5.3, and Grok-4.6 in one tester’s deliberately small run.

πŸ”¬ Pangram probed AI-text detection internals – An interpretability study of an AI-text detector showed human-versus-AI separation appearing by layer 2 and reaching perfect classification by layer 24, plus an emergent knack for guessing which model family wrote a text without any such training.

πŸ“ Jie Tang reframed scaling laws – The Z.ai founder argued that parameter count means little without data, compute, and inference in view, tracing the field from Kaplan to Chinchilla to deliberate over-training, and noting that in mixture-of-experts models more total parameters can hurt reasoning.

MLOps

πŸ“¦ Unsloth Dynamic 3.0 improved quantization – A new GGUF quantization method claimed over ten percent better top-1 accuracy at the same file size, using a calibration set refined for agentic coding and multilingual use. It drew over five million downloads in five days.

Programming

πŸ”₯ Mojo became fully open source – Modular open-sourced the Mojo compiler and toolchain under Apache 2.0 with LLVM exceptions, after four years of an open community but a closed compiler. Mojo aims to be as easy to write as Python but as fast as C, built to get full performance out of GPUs and AI chips.

πŸ“ CLAUDE.md kept growing without bound – A study of nearly 248,000 instruction lifetimes found agentic coding prompts more than tripled over time because deleting a rule meant recalling why it was added. Prompt comments that recorded intent removed most of the bloat.

Reinforcement Learning

⚑ Agent Lightning harnessed agentic RL – A Microsoft framework let the deploy-time agent harness, not the trainer, own the environment loop during reinforcement learning, in about 3,500 lines of code. With only 6K examples it lifted a 9B model on SWE-bench Verified from 41.8 to 56.4 percent.

Robotic, World AI

πŸ€– GEN-1.5 learned tasks in seconds – Generalist AI’s robot foundation model learned new manipulation tasks from a single 3-to-12-second demonstration with no training, generalizing to unseen tools and chaining behaviors it never saw.

  • One-shot in-context prompting hit 59 percent average success across ten tasks straight from pretraining.

  • Few-shot tuning on five minutes of data pushed success to 83 percent.

  • Demonstrations recorded in simulation transferred to the real robot despite no simulation data in pretraining.

Other topics

βž— OpenAI released Lean proof certificates – OpenAI published Lean 4 formalizations for ten results across mathematics and theoretical computer science, from sphere-packing bounds to a non-sofic group construction and new Ramsey number lower bounds, each independently checkable.

πŸ“‰ Reddit vanished from ChatGPT citations – An analysis reported Reddit fell from 15 percent of ChatGPT citations to zero, with review sites like G2 and Capterra also dropping out, as the model leaned toward docs, help centers, and established trusted sources.

Contributors This Week

Nancy Wang, Jeremy Ringard, Gabriel Olympie, Quentin Dubois, Glenn Sonna, Robert Hommes, NoΓ© Charmet, Yvann Barbot, Amine Saboni, Gabriel Duciel, Julien Millet, Mehdi Si-Mohammed, Pierre Chapuis, Victoire Cachoux