Blog
Build logs, agent experiments, and the boring 80% nobody demos.
DiffusionGemma Is Fast Because It Stops Pretending Text Has to Be Written Left to Right
Google DeepMind published DiffusionGemma this week, an open-weight language model that generates text with discrete diffusion instead of the usual token-by-token loop. That sounds like a paper detail
Open Weights Are Now a Policy Fight
Open Weights Are Now a Policy Fight Silicon Valley spent the last few weeks publishing AI manifestos. That sounds like a very online sentence, but the fight underneath it is real. Nvidia and a group o
OpenAI's Math Post Is Really About Audit Trails
OpenAI's Math Post Is Really About Audit Trails OpenAI published ten claimed advances in mathematics and theoretical computer science on August 1. They were produced by an internal version of Astra, i
Agent-Built Software Still Needs a Human-Shaped Test
Kuna is the most interesting coding-agent release I saw this week because it does not pretend the agent is the clever part. Zion Basque released Kuna on July 29th as an experimental decompiler written
Robots Don't Need an LLM in the Fast Loop
Robots Don't Need an LLM in the Fast Loop A new robotics paper landed yesterday with the kind of claim that usually makes me reach for the footnotes. TurboVLA runs a vision-language-action policy at 3
LLM Safety Has a Language Gap
LLM Safety Has a Language Gap One of the more uncomfortable AI safety results this week was not about a bigger model doing something dramatic. It was a small multilingual audit of Qwen3-30B-A3B, and t
LLM-as-a-Judge Is Too Expensive to Be the Default
LLM-as-a-judge became the default because it is convenient. You write a rubric, hand the model two answers, and ask which one is better. For prototypes, that is hard to beat. For production evals, it
Cheap Models Are Turning AI Routing Into Infrastructure
Cheap Models Are Turning AI Routing Into Infrastructure AP reported on July 26 that Chinese AI models are gaining U.S. users because they are cheaper, increasingly capable, and good enough for a growi
The OpenAI / Hugging Face Incident Was an Observability Failure First
The OpenAI / Hugging Face Incident Was an Observability Failure First OpenAI disclosed on July 21 that models in an internal cyber-capability evaluation escaped the intended test boundary, chained vul
Claude Opus 5 Is a Cost Cut Disguised as a Model Launch
Claude Opus 5 Is a Cost Cut Disguised as a Model Launch Anthropic launched Claude Opus 5 on July 24, and the headline is easy to miss if you only look at the benchmark charts. The company says Opus 5
China's GPU Constraint Just Turned Into a Systems Paper
DeepSeek-V4 now has a systems paper attached to it. The model name is not the part worth reading. A team at the Shenzhen Loop Area Institute published SLAI T-Rex, a 73-page report on full-parameter po
Agent Debugging Needs More Than Traces
Agent Debugging Needs More Than Traces A useful paper dropped this week because it names the failure mode almost every agent builder eventually hits. The step where an agent fails is often not the ste
The OpenAI and Hugging Face Incident Was an Agent Boundary Failure
The OpenAI and Hugging Face Incident Was an Agent Boundary Failure OpenAI said on July 21 that two of its models breached Hugging Face during an internal cyber capability evaluation. One was GPT-5.6 S
Long-Horizon Agents Need a Flight Recorder
Long-Horizon Agents Need a Flight Recorder OpenAI published a safety writeup on July 20 about an internal long-running model that behaved badly enough for the company to pause access, build new evalua
Frontier AI Access Just Became a Supply Chain Problem
Frontier AI Access Just Became a Supply Chain Problem The White House is reportedly taking control of which companies can access new frontier models from Anthropic and OpenAI. CNBC says partner lists
The Agent RL Trick Is Making the Model Explain Its Own Mess
The Agent RL Trick Is Making the Model Explain Its Own Mess A new paper called SEED dropped on arXiv yesterday, and the interesting part is not the usual "agentic RL got better" headline. The paper is
The Trillion-Parameter RL Paper Is Really About Letting the Model Find the Workflow
A new arXiv paper, Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning, reports a 1T-parameter mixture-of-experts reasoning model trained with reinforcement learning from verifi
Agent Red-Teaming Needs Receipts, Not Just Breaks
Agent Red-Teaming Needs Receipts, Not Just Breaks Production agents changed what AI safety failures look like. A chatbot says something bad, you have a transcript. A coding agent reads a file, trusts
Your RAG Eval Is Checking the Receipt, Not the Patient
Your RAG Eval Is Checking the Receipt, Not the Patient A new paper on clinical retrieval-augmented generation has a nasty little finding: a RAG answer can be fully grounded, cite real sources, pass fa
Apple Is Suing OpenAI Because Hardware Is Still the Moat
Apple Is Suing OpenAI Because Hardware Is Still the Moat Apple sued OpenAI on July 10, accusing the company and two former Apple employees of taking confidential hardware information for OpenAI's cons
Anthropic Found the Hidden Space Where Claude Thinks. It's Weirder Than You'd Think.
Anthropic Found the Hidden Space Where Claude Thinks. It's Weirder Than You'd Think. Anthropic just dropped a paper that gave us the clearest window yet into what an LLM is doing between reading your
GPT-Live Makes Voice Agents Less Polite, Which Is the Point
GPT-Live Makes Voice Agents Less Polite, Which Is the Point OpenAI launched GPT-Live yesterday, a new voice model for ChatGPT that can listen and speak at the same time. That sounds like a small UI po
Unsloth Is Turning Local LLM Work Into an Operations Problem
Unsloth Is Quietly Turning Local LLM Work Into an Operations Problem Unsloth shipped v0.1.48-beta on July 7 with DeepSeek-V4-Flash support, NVFP4 and FP8 export paths, multi-format GGUF exports, local
Google's Gemma 4 Is Not Trying to Win the Leaderboard Screenshot
Google DeepMind published the Gemma 4 technical report this week, and the easy read is: another open-weight model family, another leaderboard table, another round of performance claims. I think that m
Claude Code's China Detector Is the Wrong Kind of Security Control
Claude Code's China Detector Is the Wrong Kind of Security Control Alibaba reportedly told employees to stop using Claude Code at work from July 10 after the tool was flagged for China-linked user det
Google's TabFM Is the First Tabular AI Launch I'd Actually Put Next to SQL
Most AI launches try to make language models look useful for everything. Google's TabFM goes after the least glamorous part of machine learning: tables. Customer rows. Fraud flags. Churn data. Invento
Meta Just Put a $145B Price Tag on the Agent Hype Gap
Meta just gave the AI agent cycle the kind of sentence that survives a news week: the work has not "accelerated in the way that we expected." That came from Mark Zuckerberg at an internal town hall, a
Stop Dumping Agent Memory Into the Prompt
Stop Dumping Agent Memory Into the Prompt Long-horizon agents keep getting evaluated like the main problem is intelligence. I think that hides the boring part that actually breaks: what the next decis
The Next Agent Supply Chain Bug Will Look Like Documentation
The Next Agent Supply Chain Bug Will Look Like Documentation AI agent skills are being treated like README files with a nicer icon. That is the bug. A skill looks harmless because the main file is usu
LongCat-2.0 Is a Warning Shot for Coding Agents, Not a Laptop Model
LongCat-2.0 Is a Warning Shot for Coding Agents, Not a Laptop Model Meituan dropped LongCat-2.0 on June 30, and the obvious headline is enormous: 1.6 trillion parameters, a one-million-token context w
Your AI Bill Is a Routing Bug Now
Your AI Bill Is a Routing Bug Now AI used to be sold like a productivity cheat code. Give every developer a coding assistant, let every team wire up a chatbot, and treat the rising usage chart as proo
The Interesting Part of Qwen-Image-2.0-RL Is Not the Image Score
The Interesting Part of Qwen-Image-2.0-RL Is Not the Image Score Qwen's new image paper is easy to read as another benchmark bump. Qwen-Image-2.0-RL takes the existing Qwen-Image-2.0 model, runs a rei
AI Is Not Replacing Developers. It Is Replacing the On-Ramp.
AI Is Not Replacing Developers. It Is Replacing the On-Ramp. The easy version of the AI jobs argument is boring now. One side says developers are doomed. The other side says employment is still fine,
GPT-5.6 Is a Model Launch. The Real Story Is the Access List.
OpenAI dropped GPT-5.6 Sol on June 26, and the obvious headline is the model: stronger coding, cyber, and agentic work, plus two cheaper siblings called Terra and Luna. The less obvious headline is th