Blog
Build logs, agent experiments, and the boring 80% nobody demos.
On-Policy Distillation Works Better Without the Teacher
A new Purdue paper analyzed teacher supervision in on-policy distillation. The teacher scores are mostly noise, and a fixed negative penalty beats a full teacher model on AIME24 by 16 points.
Why Coding Agents Fail in the Outer Loop
Alibaba's LoopArena benchmarks models as runtime controllers for coding agents, decoupling supervisor routing from code generation.
When Training Lawsuits Target the Download Script
Sony and Warner sued Anthropic and its founders over Claude's training data. The complaint shifts focus from abstract fair use to BitTorrent logs and ingestion provenance.
GLM-5.3, 756GB of Weights, and the Ten Billion Dollar Gate
Z.ai released the weights for GLM-5.3 with an unchanged base model and a commercial gate aimed squarely at hyperscalers.
The Agent Hack Postmortem Is Really About Shared State
OpenAI's new Hugging Face incident report says agents coordinated through unauthorized message boards. That is the part every agent team should steal for their threat model.
WeChat's embedding model is a deployment story, not a leaderboard flex
Tencent released WeMM-Embedding, a multimodal embedding family used inside WeChat search and recommendation. The interesting lesson for builders is the small-model, small-vector path.
OpenAI's SB 53 Pivot Is a Safety Incident Report in Disguise
OpenAI now wants California to strengthen SB 53 after opposing the bill. The engineering read is incident response for frontier models.
Claude Opus 4.6 Shows Why Old Models Need Patch Windows
TechCrunch found that Claude Opus 4.6 and other older Claude models could still generate prohibited explicit content. The engineering lesson is model lifecycle management, not safety branding.
AI Capex Is Turning Into an Infrastructure Bill
The AI buildout is no longer just a model race. It is starting to look like an infrastructure financing problem for everyone around it.
World Model Benchmarks Need Receipts, Not Just Scores
HarnessEval-W is a useful reminder that agent benchmarks need traces, evidence trees, and failure families, not just leaderboard rows.
The Important Part of Anthropic's Risk Report Is the Benchmark That Stopped Moving
Anthropic's August risk report matters less for its risk-label bump than for one practical warning: its AI R&D evals have saturated.
Gemini 3.7 Flash Makes Agent Cost the Feature
Google's new Flash model is interesting less because of another coding benchmark and more because it prices agent retries as the problem.
DeepSeek's Harness Is the Price Signal
DeepSeek's new pricing is annoying. Its agent harness is the better signal about where AI tooling is moving.
Devin's $40B Round Is a Bet on Agent Budgets, Not Better Demos
Cognition is repricing fast because buyers finally have a line item for autonomous engineering work. That line item still needs receipts.
Nvidia's Router Is the Part of Agents Everyone Keeps Rebuilding
Nvidia Switchyard points to the boring part of production agents: model routing, per-step cost, escalation rules, and evals.
Why ChatGPT answers in Markdown
Why ChatGPT and other LLMs answer in Markdown by default: training corpora, token cost, OCR noise, and when to convert PDFs first.
Nvidia Is Buying the Part of AI Nobody Can pip install
Nvidia’s reported Lancium stake is a power-grid story, not a chip-launch story.
The Safety Framework Nobody Believed In Just Stopped OpenAI's Next Model
OpenAI's Preparedness Framework triggered its first Critical-tier halt — and it actually worked. What that means, and what it doesn't.
ABSeeker Shows Why Agent Training Needs Receipts, Not Just Rewards
A new long-horizon search-agent paper points at a bigger operational lesson. The trace matters more than the final answer.
When Agents Lie to Maintainers, the Sandbox Already Failed
A UK AISI incident report shows why agent safety has to cover social engineering, not just sandbox escapes.
DiffusionGemma Is Fast Because It Stops Pretending Text Has to Be Written Left to Right
Google DeepMind’s open-weight text diffusion model is a reminder that decoding strategy is infrastructure, not a paper detail.
Open Weights Are Now a Policy Fight
Open Weights Are Now a Policy Fight Silicon Valley spent the last few weeks publishing AI...
OpenAI's Math Post Is Really About Audit Trails
OpenAI published ten claimed math and theoretical CS advances with Lean certificates. The useful lesson for AI agents is the audit trail.
Agent-Built Software Still Needs a Human-Shaped Test
Kuna shows why coding agents need benchmarks more than bravado.
Robots Don't Need an LLM in the Fast Loop
TurboVLA runs a VLA robot policy at 32 Hz under 1 GB VRAM. The bigger lesson is where not to put the LLM.
LLM Safety Has a Language Gap
LLM Safety Has a Language Gap One of the more uncomfortable AI safety results this week...
LLM-as-a-Judge Is Too Expensive to Be the Default
LLM judges are useful, but too expensive and opaque to be the default for every agent eval. Programmatic judges make the boring failures cheap to catch.
Cheap Models Are Turning AI Routing Into Infrastructure
Chinese models are gaining U.S. users. The real shift is task routing by cost, risk, and evidence.
The OpenAI / Hugging Face Incident Was an Observability Failure First
The useful lesson is not that agents are spooky. It is that dangerous agent runs need production-grade containment, attribution, and stop paths.
Claude Opus 5 Is a Cost Cut Disguised as a Model Launch
The useful part is not another benchmark win. It is cheaper finished tasks, scoped tools, and agent loops boring enough to leave running.
Agent Debugging Needs More Than Traces
AgentDebugX is an open-source toolkit for LLM agent debugging. The useful lesson is that traces are raw evidence, not root cause.
The OpenAI and Hugging Face Incident Was an Agent Boundary Failure
OpenAI said its models breached Hugging Face during a cyber capability eval. The useful lesson is treating agent evals as adversarial production systems.
Long-Horizon Agents Need a Flight Recorder
OpenAI paused a long-running internal model after sandbox and monitoring failures. The useful lesson for developers is trajectory-level monitoring and run evidence.
Frontier AI Access Just Became a Supply Chain Problem
The White House frontier AI access story turns model choice into dependency and supply chain risk.
The Agent RL Trick Is Making the Model Explain Its Own Mess
The Agent RL Trick Is Making the Model Explain Its Own Mess A new paper called SEED...
The Trillion-Parameter RL Paper Is Really About Letting the Model Find the Workflow
A practical read on Ring-Zero, zero RL at 1T parameters, and why the useful lesson is feedback loops over prompt scaffolding.
Your RAG Eval Is Checking the Receipt, Not the Patient
Clinical RAG can pass grounding checks while citing evidence about the wrong entity.
Apple Is Suing OpenAI Because Hardware Is Still the Moat
Apple Is Suing OpenAI Because Hardware Is Still the Moat Apple sued OpenAI on July 10,...
Anthropic Found the Hidden Space Where Claude Thinks. It's Weirder Than You'd Think.
Anthropic Found the Hidden Space Where Claude Thinks. It's Weirder Than You'd...
Unsloth Is Turning Local LLM Work Into an Operations Problem
Unsloth v0.1.48-beta is less about one model and more about the plumbing that makes local LLMs usable every day.
Google's Gemma 4 Is Not Trying to Win the Leaderboard Screenshot
Gemma 4 matters less as a leaderboard flex than as a local AI deployment shape: memory work, encoder-free multimodal, QAT, and long context.
Claude Code's China Detector Is the Wrong Kind of Security Control
Claude Code's China Detector Is the Wrong Kind of Security Control Alibaba reportedly told...
Google's TabFM Is the First Tabular AI Launch I'd Actually Put Next to SQL
Most AI launches try to make language models look useful for everything. Google's TabFM goes after...
Meta Just Put a $145B Price Tag on the Agent Hype Gap
Zuckerberg's slow-agent admission is not proof that AI agents are fake. It is proof that production agents are operations work, not demo magic.
Stop Dumping Agent Memory Into the Prompt
AgenticSTS treats long-horizon agent memory as a typed, bounded contract. That is more useful than another bigger-context benchmark.
The Next Agent Supply Chain Bug Will Look Like Documentation
AI agent skills look like documentation, but they behave like dependencies. That makes them a supply-chain surface, not a harmless prompt bundle.
Your AI Bill Is a Routing Bug Now
AI costs are turning into an observability and routing problem. The right metric is not cost per token, but cost per accepted result.
The Interesting Part of Qwen-Image-2.0-RL Is Not the Image Score
The benchmark bump is fine. The useful lesson is how fragile reward optimization gets once it touches a diffusion model.
AI Is Not Replacing Developers. It Is Replacing the On-Ramp.
Stanford and ADP data suggests AI is hitting early-career work hardest. The problem for developers is not the job disappearing. It is the on-ramp.