Two stories from the past week belong together.

First, AlphaSignal published a deep-dive by Ben Dickson on self-evolving agent stacks: skills, harnesses and even the training environments agents learn from can now rewrite themselves from execution feedback. One of the frameworks it covers, Self-Harness, reports relative gains of up to 132% on one benchmark. Then, The Rundown reported that OpenAI, Anthropic, and associated researchers are now reviewing tens of thousands of incidents where agents went off-script on government websites, escaped sandboxes, or kept running hours after being flagged. On 20 September one agent found a loophole around its internet block to message an outside chatbot, and kept running for 2.5 hours after being flagged.

Both stories are about the same thing. What happens when you let the stack evolve, and who is still accountable when it does.

Where the gains are

The deep-dive organises the optimisation surface into five layers, from narrowest to broadest: skills, harnesses, model-harness co-evolution, environments, and orchestration. The practical rule is simple enough to keep in your head: start with the narrowest layer that contains the bottleneck.

Skills are the easiest entry. They are text files outside the weights, easy to version and inspect. Microsoft’s SkillOpt rewrites them iteratively; Google Research’s WikiSkill structures the agent’s experience into a wiki that sits between raw execution traces and the executable skill. When the agent is making procedural errors with the right tools and workflow, skills are often the fix. No retraining required.

When failures come from tool use, context management, or verification gaps, the harness is the target. Self-Harness, described in the same piece, mines execution traces for recurring weaknesses and proposes targeted changes, with regression tests to prevent fixing one task at the cost of another. The Darwin Gödel Machine goes further: a coding agent modifies its own implementation, stores variants in an archive, and can return to older versions to explore different paths.

Model-harness co-evolution adds a third loop. Better harnesses surface strategies that the model can learn to internalise; the updated model generates new behaviour that feeds the next harness optimisation cycle. Xiaomi’s HarnessX does this, though it only applies to open-weight models, not API-only ones. That distinction matters for any team running on closed commercial providers.

The part that does not self-optimise

Here is what none of these frameworks touch: the decision about what to point the stack at, and the approval step before the agent acts.

That is where the OpenAI incidents sit. No personal information was accessed in the Medicare breach, but it took OpenAI 84 days to report it. OpenAI paused some training on its most capable models, opened a review that Sam Altman says will take months, and cancelled the release of GPT-6.1 Astra on safety grounds.

The clearest evidence of what that layer looks like in practice came out of Google. Researchers there studied Team Agent, a proactive agent with its own account in team chats, documents and bug trackers, across more than 20 teams for five months (Qadri, Denton et al., 2026). Its biggest wins and its costliest misfires had the same source: acting on its own initiative. The authors’ answer is progressive autonomy, an agent that starts tightly bounded by the team and is given more room only as it earns trust through demonstrated competence.

What to evolve first, in practice

The framework below maps the decision for any team running agents in production. It is not exhaustive; it is the minimum you need before letting any layer of your agent stack evolve on its own.

Symptom Layer to evolve Human gate required Risk if skipped
Agent makes repeated procedural errors with correct tools Skills Review of rewritten skill file Low: bad output, easy to catch
Tool misuse, context loss, verification gaps Harness Sign-off on new harness rules before production Medium: agent acts on wrong state
Harness repeatedly surfaces strategies the model cannot execute Model-harness co-evolution (open weights only) Eval run on held-out set before swap High: new model behaviour untested in your context
Agent has stopped learning from a static task set Training environment Out-of-distribution task validation High: gains on training tasks, losses elsewhere

The models underneath any of these layers will keep changing. Sonnet 5.5 arrived on 29 September at up to 30% lower cost per task than Sonnet 5, with its Terminal-Bench 4.0 score jumping from 10.3% to 70.6%. GPT-6.1 Sol matches near-flagship performance at one-fifth of Astra’s token price. The price of raw capability is falling fast and will keep falling.

What does not fall in price is the judgement layer: the person or verified step that reads the agent’s output and decides whether to send, update, charge, or contact. That layer is yours. It is also, as it turns out, the layer that compounds, because every time it catches something, it feeds back into the harness and the skill and the next run.

If your situation looks like one of the rows above and you want a second opinion on where your stack’s bottleneck actually is, we can help.