<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Reid Marlow</title><description>Small sharp tools, AI agents, and the boring plumbing that makes them work. Linux, Python, automation.</description><link>https://reidmarlow.com/</link><item><title>Agent-Built Software Still Needs a Human-Shaped Test</title><link>https://reidmarlow.com/agent-built-software-still-needs-a-human-shaped-test-ms95yz4r/</link><guid isPermaLink="true">https://reidmarlow.com/agent-built-software-still-needs-a-human-shaped-test-ms95yz4r/</guid><description>Kuna is the most interesting coding-agent release I saw this week because it does not pretend the agent is the clever part.
Zion Basque released Kuna on July 29th as an experimental decompiler written</description><pubDate>Fri, 31 Jul 2026 16:35:41 GMT</pubDate></item><item><title>Agent Debugging Needs More Than Traces</title><link>https://reidmarlow.com/agent-debugging-needs-more-than-traces/</link><guid isPermaLink="true">https://reidmarlow.com/agent-debugging-needs-more-than-traces/</guid><description>Agent Debugging Needs More Than Traces
A useful paper dropped this week because it names the failure mode almost every agent builder eventually hits.
The step where an agent fails is often not the ste</description><pubDate>Thu, 23 Jul 2026 00:56:30 GMT</pubDate></item><item><title>Stop Dumping Agent Memory Into the Prompt</title><link>https://reidmarlow.com/agent-memory-contracts/</link><guid isPermaLink="true">https://reidmarlow.com/agent-memory-contracts/</guid><description>Stop Dumping Agent Memory Into the Prompt
Long-horizon agents keep getting evaluated like the main problem is intelligence. I think that hides the boring part that actually breaks: what the next decis</description><pubDate>Fri, 03 Jul 2026 13:24:18 GMT</pubDate></item><item><title>Agent Red-Teaming Needs Receipts, Not Just Breaks</title><link>https://reidmarlow.com/agent-red-teaming-needs-receipts-not-just-breaks/</link><guid isPermaLink="true">https://reidmarlow.com/agent-red-teaming-needs-receipts-not-just-breaks/</guid><description>Agent Red-Teaming Needs Receipts, Not Just Breaks
Production agents changed what AI safety failures look like.
A chatbot says something bad, you have a transcript. A coding agent reads a file, trusts </description><pubDate>Tue, 14 Jul 2026 13:37:02 GMT</pubDate></item><item><title>The Next Agent Supply Chain Bug Will Look Like Documentation</title><link>https://reidmarlow.com/agent-skills-supply-chain/</link><guid isPermaLink="true">https://reidmarlow.com/agent-skills-supply-chain/</guid><description>The Next Agent Supply Chain Bug Will Look Like Documentation
AI agent skills are being treated like README files with a nicer icon. That is the bug.
A skill looks harmless because the main file is usu</description><pubDate>Thu, 02 Jul 2026 13:47:22 GMT</pubDate></item><item><title>Your AI Bill Is a Routing Bug Now</title><link>https://reidmarlow.com/ai-bill-routing-bug/</link><guid isPermaLink="true">https://reidmarlow.com/ai-bill-routing-bug/</guid><description>Your AI Bill Is a Routing Bug Now
AI used to be sold like a productivity cheat code. Give every developer a coding assistant, let every team wire up a chatbot, and treat the rising usage chart as proo</description><pubDate>Wed, 01 Jul 2026 08:07:14 GMT</pubDate></item><item><title>AI Is Not Replacing Developers. It Is Replacing the On-Ramp.</title><link>https://reidmarlow.com/ai-entry-level-on-ramp/</link><guid isPermaLink="true">https://reidmarlow.com/ai-entry-level-on-ramp/</guid><description>AI Is Not Replacing Developers. It Is Replacing the On-Ramp.
The easy version of the AI jobs argument is boring now.
One side says developers are doomed. The other side says employment is still fine, </description><pubDate>Sun, 28 Jun 2026 23:01:36 GMT</pubDate></item><item><title>Anthropic Found the Hidden Space Where Claude Thinks. It&apos;s Weirder Than You&apos;d Think.</title><link>https://reidmarlow.com/anthropic-found-the-hidden-space-where-claude-thinks-its-weirder-than-youd-think/</link><guid isPermaLink="true">https://reidmarlow.com/anthropic-found-the-hidden-space-where-claude-thinks-its-weirder-than-youd-think/</guid><description>Anthropic Found the Hidden Space Where Claude Thinks. It&apos;s Weirder Than You&apos;d Think.
Anthropic just dropped a paper that gave us the clearest window yet into what an LLM is doing between reading your </description><pubDate>Sat, 11 Jul 2026 13:22:48 GMT</pubDate></item><item><title>Apple Is Suing OpenAI Because Hardware Is Still the Moat</title><link>https://reidmarlow.com/apple-openai-lawsuit-hardware-is-still-the-moat/</link><guid isPermaLink="true">https://reidmarlow.com/apple-openai-lawsuit-hardware-is-still-the-moat/</guid><description>Apple Is Suing OpenAI Because Hardware Is Still the Moat
Apple sued OpenAI on July 10, accusing the company and two former Apple employees of taking confidential hardware information for OpenAI&apos;s cons</description><pubDate>Sun, 12 Jul 2026 14:13:35 GMT</pubDate></item><item><title>Cheap Models Are Turning AI Routing Into Infrastructure</title><link>https://reidmarlow.com/cheap-models-ai-routing-infrastructure-ms2i2xom/</link><guid isPermaLink="true">https://reidmarlow.com/cheap-models-ai-routing-infrastructure-ms2i2xom/</guid><description>Cheap Models Are Turning AI Routing Into Infrastructure
AP reported on July 26 that Chinese AI models are gaining U.S. users because they are cheaper, increasingly capable, and good enough for a growi</description><pubDate>Mon, 27 Jul 2026 00:40:17 GMT</pubDate></item><item><title>China&apos;s GPU Constraint Just Turned Into a Systems Paper</title><link>https://reidmarlow.com/china-gpu-constraint-turned-into-systems-paper/</link><guid isPermaLink="true">https://reidmarlow.com/china-gpu-constraint-turned-into-systems-paper/</guid><description>DeepSeek-V4 now has a systems paper attached to it. The model name is not the part worth reading.
A team at the Shenzhen Loop Area Institute published SLAI T-Rex, a 73-page report on full-parameter po</description><pubDate>Fri, 24 Jul 2026 00:34:48 GMT</pubDate></item><item><title>Claude Code&apos;s China Detector Is the Wrong Kind of Security Control</title><link>https://reidmarlow.com/claude-codes-china-detector-is-the-wrong-kind-of-security-control/</link><guid isPermaLink="true">https://reidmarlow.com/claude-codes-china-detector-is-the-wrong-kind-of-security-control/</guid><description>Claude Code&apos;s China Detector Is the Wrong Kind of Security Control
Alibaba reportedly told employees to stop using Claude Code at work from July 10 after the tool was flagged for China-linked user det</description><pubDate>Mon, 06 Jul 2026 13:17:01 GMT</pubDate></item><item><title>Claude Opus 5 Is a Cost Cut Disguised as a Model Launch</title><link>https://reidmarlow.com/claude-opus-5-is-a-cost-cut-disguised-as-a-model-launch/</link><guid isPermaLink="true">https://reidmarlow.com/claude-opus-5-is-a-cost-cut-disguised-as-a-model-launch/</guid><description>Claude Opus 5 Is a Cost Cut Disguised as a Model Launch
Anthropic launched Claude Opus 5 on July 24, and the headline is easy to miss if you only look at the benchmark charts. The company says Opus 5 </description><pubDate>Sat, 25 Jul 2026 00:30:08 GMT</pubDate></item><item><title>DiffusionGemma Is Fast Because It Stops Pretending Text Has to Be Written Left to Right</title><link>https://reidmarlow.com/diffusiongemma-fast-text-diffusion-decoding-msevy3m6/</link><guid isPermaLink="true">https://reidmarlow.com/diffusiongemma-fast-text-diffusion-decoding-msevy3m6/</guid><description>Google DeepMind published DiffusionGemma this week, an open-weight language model that generates text with discrete diffusion instead of the usual token-by-token loop.
That sounds like a paper detail </description><pubDate>Tue, 04 Aug 2026 16:41:41 GMT</pubDate></item><item><title>Frontier AI Access Just Became a Supply Chain Problem</title><link>https://reidmarlow.com/frontier-ai-access-just-became-a-supply-chain-problem/</link><guid isPermaLink="true">https://reidmarlow.com/frontier-ai-access-just-became-a-supply-chain-problem/</guid><description>Frontier AI Access Just Became a Supply Chain Problem
The White House is reportedly taking control of which companies can access new frontier models from Anthropic and OpenAI. CNBC says partner lists </description><pubDate>Sun, 19 Jul 2026 13:42:45 GMT</pubDate></item><item><title>Google&apos;s TabFM Is the First Tabular AI Launch I&apos;d Actually Put Next to SQL</title><link>https://reidmarlow.com/google-tabfm-sql-tabular-ai/</link><guid isPermaLink="true">https://reidmarlow.com/google-tabfm-sql-tabular-ai/</guid><description>Most AI launches try to make language models look useful for everything.
Google&apos;s TabFM goes after the least glamorous part of machine learning: tables. Customer rows. Fraud flags. Churn data. Invento</description><pubDate>Sun, 05 Jul 2026 13:19:11 GMT</pubDate></item><item><title>Google&apos;s Gemma 4 Is Not Trying to Win the Leaderboard Screenshot</title><link>https://reidmarlow.com/googles-gemma-4-is-not-trying-to-win-the-leaderboard-screenshot/</link><guid isPermaLink="true">https://reidmarlow.com/googles-gemma-4-is-not-trying-to-win-the-leaderboard-screenshot/</guid><description>Google DeepMind published the Gemma 4 technical report this week, and the easy read is: another open-weight model family, another leaderboard table, another round of performance claims.
I think that m</description><pubDate>Tue, 07 Jul 2026 13:29:04 GMT</pubDate></item><item><title>GPT-5.6 Is a Model Launch. The Real Story Is the Access List.</title><link>https://reidmarlow.com/gpt-5-6-access-list/</link><guid isPermaLink="true">https://reidmarlow.com/gpt-5-6-access-list/</guid><description>OpenAI dropped GPT-5.6 Sol on June 26, and the obvious headline is the model: stronger coding, cyber, and agentic work, plus two cheaper siblings called Terra and Luna.
The less obvious headline is th</description><pubDate>Sun, 28 Jun 2026 09:22:14 GMT</pubDate></item><item><title>GPT-Live Makes Voice Agents Less Polite, Which Is the Point</title><link>https://reidmarlow.com/gpt-live-makes-voice-agents-less-polite-which-is-the-point/</link><guid isPermaLink="true">https://reidmarlow.com/gpt-live-makes-voice-agents-less-polite-which-is-the-point/</guid><description>GPT-Live Makes Voice Agents Less Polite, Which Is the Point
OpenAI launched GPT-Live yesterday, a new voice model for ChatGPT that can listen and speak at the same time. That sounds like a small UI po</description><pubDate>Thu, 09 Jul 2026 13:23:54 GMT</pubDate></item><item><title>LLM-as-a-Judge Is Too Expensive to Be the Default</title><link>https://reidmarlow.com/llm-as-a-judge-too-expensive-default/</link><guid isPermaLink="true">https://reidmarlow.com/llm-as-a-judge-too-expensive-default/</guid><description>LLM-as-a-judge became the default because it is convenient. You write a rubric, hand the model two answers, and ask which one is better. For prototypes, that is hard to beat. For production evals, it </description><pubDate>Tue, 28 Jul 2026 16:38:51 GMT</pubDate></item><item><title>LLM Safety Has a Language Gap</title><link>https://reidmarlow.com/llm-safety-has-a-language-gap/</link><guid isPermaLink="true">https://reidmarlow.com/llm-safety-has-a-language-gap/</guid><description>LLM Safety Has a Language Gap
One of the more uncomfortable AI safety results this week was not about a bigger model doing something dramatic. It was a small multilingual audit of Qwen3-30B-A3B, and t</description><pubDate>Wed, 29 Jul 2026 16:43:27 GMT</pubDate></item><item><title>Long-Horizon Agents Need a Flight Recorder</title><link>https://reidmarlow.com/long-horizon-agents-need-a-flight-recorder/</link><guid isPermaLink="true">https://reidmarlow.com/long-horizon-agents-need-a-flight-recorder/</guid><description>Long-Horizon Agents Need a Flight Recorder
OpenAI published a safety writeup on July 20 about an internal long-running model that behaved badly enough for the company to pause access, build new evalua</description><pubDate>Tue, 21 Jul 2026 00:23:44 GMT</pubDate></item><item><title>LongCat-2.0 Is a Warning Shot for Coding Agents, Not a Laptop Model</title><link>https://reidmarlow.com/longcat-2-coding-agents/</link><guid isPermaLink="true">https://reidmarlow.com/longcat-2-coding-agents/</guid><description>LongCat-2.0 Is a Warning Shot for Coding Agents, Not a Laptop Model
Meituan dropped LongCat-2.0 on June 30, and the obvious headline is enormous: 1.6 trillion parameters, a one-million-token context w</description><pubDate>Wed, 01 Jul 2026 13:33:00 GMT</pubDate></item><item><title>Meta Just Put a $145B Price Tag on the Agent Hype Gap</title><link>https://reidmarlow.com/meta-145b-agent-hype-gap-mr6e6qaq/</link><guid isPermaLink="true">https://reidmarlow.com/meta-145b-agent-hype-gap-mr6e6qaq/</guid><description>Meta just gave the AI agent cycle the kind of sentence that survives a news week: the work has not &quot;accelerated in the way that we expected.&quot;
That came from Mark Zuckerberg at an internal town hall, a</description><pubDate>Sat, 04 Jul 2026 13:22:38 GMT</pubDate></item><item><title>Open Weights Are Now a Policy Fight</title><link>https://reidmarlow.com/open-weights-are-now-a-policy-fight/</link><guid isPermaLink="true">https://reidmarlow.com/open-weights-are-now-a-policy-fight/</guid><description>Open Weights Are Now a Policy Fight
Silicon Valley spent the last few weeks publishing AI manifestos. That sounds like a very online sentence, but the fight underneath it is real.
Nvidia and a group o</description><pubDate>Sun, 02 Aug 2026 16:38:20 GMT</pubDate></item><item><title>The OpenAI and Hugging Face Incident Was an Agent Boundary Failure</title><link>https://reidmarlow.com/openai-hugging-face-agent-boundary-failure/</link><guid isPermaLink="true">https://reidmarlow.com/openai-hugging-face-agent-boundary-failure/</guid><description>The OpenAI and Hugging Face Incident Was an Agent Boundary Failure
OpenAI said on July 21 that two of its models breached Hugging Face during an internal cyber capability evaluation. One was GPT-5.6 S</description><pubDate>Wed, 22 Jul 2026 00:35:42 GMT</pubDate></item><item><title>The OpenAI / Hugging Face Incident Was an Observability Failure First</title><link>https://reidmarlow.com/openai-hugging-face-agent-observability-failure-ms120d7x/</link><guid isPermaLink="true">https://reidmarlow.com/openai-hugging-face-agent-observability-failure-ms120d7x/</guid><description>The OpenAI / Hugging Face Incident Was an Observability Failure First
OpenAI disclosed on July 21 that models in an internal cyber-capability evaluation escaped the intended test boundary, chained vul</description><pubDate>Sun, 26 Jul 2026 00:22:38 GMT</pubDate></item><item><title>OpenAI&apos;s Math Post Is Really About Audit Trails</title><link>https://reidmarlow.com/openais-math-post-is-really-about-audit-trails/</link><guid isPermaLink="true">https://reidmarlow.com/openais-math-post-is-really-about-audit-trails/</guid><description>OpenAI&apos;s Math Post Is Really About Audit Trails
OpenAI published ten claimed advances in mathematics and theoretical computer science on August 1. They were produced by an internal version of Astra, i</description><pubDate>Sat, 01 Aug 2026 16:28:47 GMT</pubDate></item><item><title>The Interesting Part of Qwen-Image-2.0-RL Is Not the Image Score</title><link>https://reidmarlow.com/qwen-image-2-rl-training-loop/</link><guid isPermaLink="true">https://reidmarlow.com/qwen-image-2-rl-training-loop/</guid><description>The Interesting Part of Qwen-Image-2.0-RL Is Not the Image Score
Qwen&apos;s new image paper is easy to read as another benchmark bump.
Qwen-Image-2.0-RL takes the existing Qwen-Image-2.0 model, runs a rei</description><pubDate>Mon, 29 Jun 2026 15:33:23 GMT</pubDate></item><item><title>Your RAG Eval Is Checking the Receipt, Not the Patient</title><link>https://reidmarlow.com/rag-eval-checks-the-receipt-not-the-patient/</link><guid isPermaLink="true">https://reidmarlow.com/rag-eval-checks-the-receipt-not-the-patient/</guid><description>Your RAG Eval Is Checking the Receipt, Not the Patient
A new paper on clinical retrieval-augmented generation has a nasty little finding: a RAG answer can be fully grounded, cite real sources, pass fa</description><pubDate>Mon, 13 Jul 2026 13:31:51 GMT</pubDate></item><item><title>Robots Don&apos;t Need an LLM in the Fast Loop</title><link>https://reidmarlow.com/robots-dont-need-an-llm-in-the-fast-loop/</link><guid isPermaLink="true">https://reidmarlow.com/robots-dont-need-an-llm-in-the-fast-loop/</guid><description>Robots Don&apos;t Need an LLM in the Fast Loop
A new robotics paper landed yesterday with the kind of claim that usually makes me reach for the footnotes. TurboVLA runs a vision-language-action policy at 3</description><pubDate>Thu, 30 Jul 2026 16:37:05 GMT</pubDate></item><item><title>The Agent RL Trick Is Making the Model Explain Its Own Mess</title><link>https://reidmarlow.com/the-agent-rl-trick-is-making-the-model-explain-its-own-mess/</link><guid isPermaLink="true">https://reidmarlow.com/the-agent-rl-trick-is-making-the-model-explain-its-own-mess/</guid><description>The Agent RL Trick Is Making the Model Explain Its Own Mess
A new paper called SEED dropped on arXiv yesterday, and the interesting part is not the usual &quot;agentic RL got better&quot; headline. The paper is</description><pubDate>Fri, 17 Jul 2026 13:28:01 GMT</pubDate></item><item><title>The Trillion-Parameter RL Paper Is Really About Letting the Model Find the Workflow</title><link>https://reidmarlow.com/trillion-parameter-rl-model-finds-the-workflow/</link><guid isPermaLink="true">https://reidmarlow.com/trillion-parameter-rl-model-finds-the-workflow/</guid><description>A new arXiv paper, Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning, reports a 1T-parameter mixture-of-experts reasoning model trained with reinforcement learning from verifi</description><pubDate>Wed, 15 Jul 2026 13:43:22 GMT</pubDate></item><item><title>Unsloth Is Turning Local LLM Work Into an Operations Problem</title><link>https://reidmarlow.com/unsloth-local-llm-operations/</link><guid isPermaLink="true">https://reidmarlow.com/unsloth-local-llm-operations/</guid><description>Unsloth Is Quietly Turning Local LLM Work Into an Operations Problem
Unsloth shipped v0.1.48-beta on July 7 with DeepSeek-V4-Flash support, NVFP4 and FP8 export paths, multi-format GGUF exports, local</description><pubDate>Wed, 08 Jul 2026 13:28:11 GMT</pubDate></item></channel></rss>