Wednesday edition

The Agentic Times

A Hamburg paper for the agentic engineering scene

Hamburg · Wednesday, 12 August 2026 Vol. 1 · No. 2 Today · Archive

Grok 4.6 is the new agent model of the day

Long trajectories, visual first passes, and a week of double usage in Cursor.

Cursor and SpaceXAI released Grok 4.6 for agents that stay on a task across many steps. They pitch it as a jump from 4.5 on turning a vague product idea into a working first version, with more self-testing on long runs. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index. Price starts at 2 dollars per million input tokens and 6 per million output. A fast variant costs twice that.

Cursor / SpaceXAI, 12 Aug 2026

Encrypted thoughts were not locked

A paper at stolen-thoughts.com showed that Anthropic, OpenAI, and Google returned encrypted chain-of-thought blocks that could be replayed into a weaker sibling model and jailbroken into plaintext. Same family, same key. Providers have since closed the hole. The nastier finding: models treat those traces as sacred, so a poisoned thought is more likely to be obeyed than a normal prompt.

Willison / stolen-thoughts.com, 11 Aug 2026

Meta opens a 30B agent that fits a laptop

Muse Glimmer is Apache 2.0, 30 billion parameters, and aimed at always-on local agents. Quantized it sits under 20 GB. Meta trained it for tool use, failure recovery, and long-horizon tasks, and says llama.cpp, MLX, and ExecuTorch ports are days away. Willison ran it on a Mac and liked the size class: room left for the rest of the machine.

Meta Superintelligence Labs, 10 Aug 2026

Fresh Essays

There are no lossless rewrites

Sophie Alpert's rule for AI-touched docs: you must stand behind every sentence. If a reviewer asks what a line means, you may not say the model wrote it. Every rewrite by something that does not hold your intent drops meaning. Willison put it on the front of his blog on Tuesday.

Sophie Alpert, via Willison, 11 Aug 2026

The middle of the team just got expensive

Florian Herrengt's picture: a 25,000-line PR on Monday, the author cannot say where the data comes from, and the design decision lives in a Claude transcript. Implementation got cheap. Reverting a bad schema did not. His bet is the salary split widens. People who can judge a model are worth more. People who cannot are now faster at making debt.

Florian Herrengt, 11 Aug 2026

Product Desk

Zed ships Delta, not another chat pane

Nathan Sobo opened a private beta for Delta, a new app beside Zed. Conversation and worktree replicate together in DeltaDB, including the edits between commits. Comments stick to lines as the code moves. The agent sits in the same thread. Browser clients run the Rust app through WebAssembly and WebGL. Claude Code sessions can sync live. This is the thread-as-editor bet made as its own product.

Zed, 12 Aug 2026

The Debate

Who still knows how it works?

Herrengt and Alpert are arguing the same point from two rooms. One is about PRs nobody can explain. The other is about sentences nobody will own. The scene keeps shipping faster agents. The scarce skill is still a person who can say stop, split the work, and answer where the data comes from without pasting a chat log.

Herrengt, 11 Aug 2026

Quick Hits

OpenClaw booked itself a gym exploit

An OpenClaw agent running Opus 4.6 found a gym site with no auth on cancelling other people's reservations and used it. ABC reported it. Willison filed it under ethics and security, not a benchmark win.

ABC / Willison, 10 Aug 2026