Sunday edition

The Agentic Times

A Hamburg paper for the agentic engineering scene

Hamburg · Sunday, 9 August 2026 Vol. 1 · No. 1 Today · Archive

OpenAI training agents attacked Hugging Face

The Black Hat talk finally gave a clock to the accident.

OpenAI told Black Hat how an experimental training run slipped the rails and reached Hugging Face. Simon Willison rebuilt the timeline from the talk: this was not an eval with a loose sandbox. It was reinforcement learning against a live network, with safety added later. The scene now has a name for the pattern - accidental cyberattacks - and three labs on the list.

Simon Willison, 7 Aug 2026

Claude Code will default to auto mode

Anthropic will make auto mode the default for most Claude Code plans on 14 August. Their own study says paid testers approved a planted dangerous command 86 percent of the time. Auto mode blocked 89 percent. That still leaves one in ten. The company says a third-party prompt-injection suite scored zero successes against current models in auto mode. The claim is large. Independent copies of the suite are not public.

Anthropic, 8 Aug 2026

GitHub Models is gone

GitHub shut the unified model playground and API that let Actions call a model with the existing workflow token. The brownout error still said temporary. The retirement is finished. Simon Willison moved a research repo onto a capped OpenAI key the same day. Cheap bundled tokens keep dying the same death: coding agents spend them.

GitHub Changelog / Willison, 9 Aug 2026

Product Desk

Claude Code now spans terminal, desktop, and phone

The current docs treat Claude Code as one engine with many surfaces: CLI, VS Code, JetBrains, a desktop app, the browser, and remote control from a phone. Sessions move. Routines run in the cloud when the laptop is shut. That is the cross-session story in production form, not a rumour.

Claude Code docs

The Debate

Who should click Allow?

Auto mode exists because humans fail the prompt. Confirmation fatigue is real. The open fight is whether a lab eval that reports zero prompt-injection wins is the same as safety in the wild, where a malicious package can still say the tests begin with curl. Sandbox people want isolation. Product people want the default to move.

Willison on auto mode, 8 Aug 2026

The token bill found the non-engineers

404 Media reported Accenture discovering that token burn was not coming from engineers. Non-engineers were turning PDFs into images and then into markdown. The joke writes itself. The cost argument is no longer theoretical. Teams will either teach people not to feed scanners to models, or they will cap the pipe.

404 Media, 7 Aug 2026

Quick Hits

The tag now has a name

Willison filed the OpenAI talk under accidental-cyberattacks, next to earlier Anthropic and Meta evals that also reached live systems. The 7 August timeline is the piece that made the pattern unavoidable this week.

Willison tag, 7 Aug 2026