Reid Marlow
$ ls -t posts/

Blog

Build logs, agent experiments, and the boring 80% nobody demos.

DiffusionGemma Is Fast Because It Stops Pretending Text Has to Be Written Left to Right

Google DeepMind published DiffusionGemma this week, an open-weight language model that generates text with discrete diffusion instead of the usual token-by-token loop. That sounds like a paper detail

Open Weights Are Now a Policy Fight

Open Weights Are Now a Policy Fight Silicon Valley spent the last few weeks publishing AI manifestos. That sounds like a very online sentence, but the fight underneath it is real. Nvidia and a group o

OpenAI's Math Post Is Really About Audit Trails

OpenAI's Math Post Is Really About Audit Trails OpenAI published ten claimed advances in mathematics and theoretical computer science on August 1. They were produced by an internal version of Astra, i

Agent-Built Software Still Needs a Human-Shaped Test

Kuna is the most interesting coding-agent release I saw this week because it does not pretend the agent is the clever part. Zion Basque released Kuna on July 29th as an experimental decompiler written

Robots Don't Need an LLM in the Fast Loop

Robots Don't Need an LLM in the Fast Loop A new robotics paper landed yesterday with the kind of claim that usually makes me reach for the footnotes. TurboVLA runs a vision-language-action policy at 3

LLM Safety Has a Language Gap

LLM Safety Has a Language Gap One of the more uncomfortable AI safety results this week was not about a bigger model doing something dramatic. It was a small multilingual audit of Qwen3-30B-A3B, and t

LLM-as-a-Judge Is Too Expensive to Be the Default

LLM-as-a-judge became the default because it is convenient. You write a rubric, hand the model two answers, and ask which one is better. For prototypes, that is hard to beat. For production evals, it

Cheap Models Are Turning AI Routing Into Infrastructure

Cheap Models Are Turning AI Routing Into Infrastructure AP reported on July 26 that Chinese AI models are gaining U.S. users because they are cheaper, increasingly capable, and good enough for a growi

The OpenAI / Hugging Face Incident Was an Observability Failure First

The OpenAI / Hugging Face Incident Was an Observability Failure First OpenAI disclosed on July 21 that models in an internal cyber-capability evaluation escaped the intended test boundary, chained vul

Claude Opus 5 Is a Cost Cut Disguised as a Model Launch

Claude Opus 5 Is a Cost Cut Disguised as a Model Launch Anthropic launched Claude Opus 5 on July 24, and the headline is easy to miss if you only look at the benchmark charts. The company says Opus 5

China's GPU Constraint Just Turned Into a Systems Paper

DeepSeek-V4 now has a systems paper attached to it. The model name is not the part worth reading. A team at the Shenzhen Loop Area Institute published SLAI T-Rex, a 73-page report on full-parameter po

Agent Debugging Needs More Than Traces

Agent Debugging Needs More Than Traces A useful paper dropped this week because it names the failure mode almost every agent builder eventually hits. The step where an agent fails is often not the ste

The OpenAI and Hugging Face Incident Was an Agent Boundary Failure

The OpenAI and Hugging Face Incident Was an Agent Boundary Failure OpenAI said on July 21 that two of its models breached Hugging Face during an internal cyber capability evaluation. One was GPT-5.6 S

Long-Horizon Agents Need a Flight Recorder

Long-Horizon Agents Need a Flight Recorder OpenAI published a safety writeup on July 20 about an internal long-running model that behaved badly enough for the company to pause access, build new evalua

Frontier AI Access Just Became a Supply Chain Problem

Frontier AI Access Just Became a Supply Chain Problem The White House is reportedly taking control of which companies can access new frontier models from Anthropic and OpenAI. CNBC says partner lists

The Agent RL Trick Is Making the Model Explain Its Own Mess

The Agent RL Trick Is Making the Model Explain Its Own Mess A new paper called SEED dropped on arXiv yesterday, and the interesting part is not the usual "agentic RL got better" headline. The paper is

The Trillion-Parameter RL Paper Is Really About Letting the Model Find the Workflow

A new arXiv paper, Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning, reports a 1T-parameter mixture-of-experts reasoning model trained with reinforcement learning from verifi

Agent Red-Teaming Needs Receipts, Not Just Breaks

Agent Red-Teaming Needs Receipts, Not Just Breaks Production agents changed what AI safety failures look like. A chatbot says something bad, you have a transcript. A coding agent reads a file, trusts

Your RAG Eval Is Checking the Receipt, Not the Patient

Your RAG Eval Is Checking the Receipt, Not the Patient A new paper on clinical retrieval-augmented generation has a nasty little finding: a RAG answer can be fully grounded, cite real sources, pass fa

Apple Is Suing OpenAI Because Hardware Is Still the Moat

Apple Is Suing OpenAI Because Hardware Is Still the Moat Apple sued OpenAI on July 10, accusing the company and two former Apple employees of taking confidential hardware information for OpenAI's cons

Anthropic Found the Hidden Space Where Claude Thinks. It's Weirder Than You'd Think.

Anthropic Found the Hidden Space Where Claude Thinks. It's Weirder Than You'd Think. Anthropic just dropped a paper that gave us the clearest window yet into what an LLM is doing between reading your

GPT-Live Makes Voice Agents Less Polite, Which Is the Point

GPT-Live Makes Voice Agents Less Polite, Which Is the Point OpenAI launched GPT-Live yesterday, a new voice model for ChatGPT that can listen and speak at the same time. That sounds like a small UI po

Unsloth Is Turning Local LLM Work Into an Operations Problem

Unsloth Is Quietly Turning Local LLM Work Into an Operations Problem Unsloth shipped v0.1.48-beta on July 7 with DeepSeek-V4-Flash support, NVFP4 and FP8 export paths, multi-format GGUF exports, local

Google's Gemma 4 Is Not Trying to Win the Leaderboard Screenshot

Google DeepMind published the Gemma 4 technical report this week, and the easy read is: another open-weight model family, another leaderboard table, another round of performance claims. I think that m

Claude Code's China Detector Is the Wrong Kind of Security Control

Claude Code's China Detector Is the Wrong Kind of Security Control Alibaba reportedly told employees to stop using Claude Code at work from July 10 after the tool was flagged for China-linked user det

Google's TabFM Is the First Tabular AI Launch I'd Actually Put Next to SQL

Most AI launches try to make language models look useful for everything. Google's TabFM goes after the least glamorous part of machine learning: tables. Customer rows. Fraud flags. Churn data. Invento

Meta Just Put a $145B Price Tag on the Agent Hype Gap

Meta just gave the AI agent cycle the kind of sentence that survives a news week: the work has not "accelerated in the way that we expected." That came from Mark Zuckerberg at an internal town hall, a

Stop Dumping Agent Memory Into the Prompt

Stop Dumping Agent Memory Into the Prompt Long-horizon agents keep getting evaluated like the main problem is intelligence. I think that hides the boring part that actually breaks: what the next decis

The Next Agent Supply Chain Bug Will Look Like Documentation

The Next Agent Supply Chain Bug Will Look Like Documentation AI agent skills are being treated like README files with a nicer icon. That is the bug. A skill looks harmless because the main file is usu

LongCat-2.0 Is a Warning Shot for Coding Agents, Not a Laptop Model

LongCat-2.0 Is a Warning Shot for Coding Agents, Not a Laptop Model Meituan dropped LongCat-2.0 on June 30, and the obvious headline is enormous: 1.6 trillion parameters, a one-million-token context w

Your AI Bill Is a Routing Bug Now

Your AI Bill Is a Routing Bug Now AI used to be sold like a productivity cheat code. Give every developer a coding assistant, let every team wire up a chatbot, and treat the rising usage chart as proo

The Interesting Part of Qwen-Image-2.0-RL Is Not the Image Score

The Interesting Part of Qwen-Image-2.0-RL Is Not the Image Score Qwen's new image paper is easy to read as another benchmark bump. Qwen-Image-2.0-RL takes the existing Qwen-Image-2.0 model, runs a rei

AI Is Not Replacing Developers. It Is Replacing the On-Ramp.

AI Is Not Replacing Developers. It Is Replacing the On-Ramp. The easy version of the AI jobs argument is boring now. One side says developers are doomed. The other side says employment is still fine,

GPT-5.6 Is a Model Launch. The Real Story Is the Access List.

OpenAI dropped GPT-5.6 Sol on June 26, and the obvious headline is the model: stronger coding, cyber, and agentic work, plus two cheaper siblings called Terra and Luna. The less obvious headline is th