Writing
September 22, 2026 · 7 min read

Recursive self-improvement of agent harnesses needs a regularizer

A new paper, RRSI, shows that letting an agent rewrite its own harness overfits to the training benchmark unless you constrain how it proposes and selects edits — the fix combines an annealed edit budget, novelty-seeking search, and a critic-plus-pruner selector.

agent-harnessself-improvementregularizationllm-agentsgeneralization

I keep coming back to the same operational question when a team asks me to automate improvements to an agent's harness — the prompts, tool wiring, memory, and control flow around a frozen model: will the gains still be there once the agent leaves the benchmark it was tuned on? A paper posted this week, RRSI: Regularized Recursive Self-Improvement of Agent Harnesses (Xia, Han, Wang, et al.), gives that question a name and a fix. It's worth reading closely, because the failure mode it describes is one I'd bet most people running harness-optimization loops in production haven't measured yet — they just haven't looked at the out-of-distribution numbers.

The setup: harnesses that edit themselves

A harness is everything wrapped around the backbone model — system prompts, retrieval and memory logic, tool definitions, control flow for retries and verification. Swap the harness and a fixed model's measured capability can swing dramatically, which is why an entire line of work now automates harness editing: propose a component-wise change, run it against a benchmark, keep it if it scores better, repeat. Do this iteratively and you get something like recursive self-improvement (RSI) at the systems level, not the weights level.

The paper's core observation is that this loop behaves exactly like an unregularized optimizer on any other objective: given enough iterations against a fixed training split, it memorizes that split. The harness accumulates edits that are really overfitting artifacts — brittle heuristics, benchmark-specific string matching, special-cased control flow — dressed up as general improvements. In-distribution scores climb; the same harness evaluated on benchmarks it never saw during evolution can lose most or all of that gain.

Three regularizers, applied where the loop actually breaks

RRSI doesn't touch the model. It regularizes the two places the evolutionary loop makes decisions: proposing candidate edits, and selecting which ones survive.

On the proposal side, two mechanisms:

  • A temporally annealed edit budget. Early in evolution, a candidate can bundle many simultaneous changes to the harness; as evolution proceeds, the budget shrinks, forcing later edits to be small and targeted rather than large compound rewrites. This is the same intuition as a learning-rate schedule — take bigger steps when you're far from a good region, smaller ones as you converge, so you don't blow past a genuinely general improvement in favor of a complex one that only works because it matches training-set quirks.
  • Novelty-seeking proposals. The proposer is biased toward unexplored regions of edit-history space rather than re-treading variations of what already scored well. This pushes the search away from piling up near-duplicate, overfit variants of the same trick.

On the selection side, a two-stage gate:

  • A critic screens proposals for being benchmark-specific rather than general — the part of the loop that's supposed to catch "this only works because of how this one eval is formatted."
  • A pruner removes survivors that are too small to matter, too expensive in tokens or latency to justify their gain, or that have simply stopped being useful as the harness evolves around them.

The pruner detail is the one I'd flag to anyone running this kind of loop internally: harness edits accrue cruft the same way codebases do. An edit that helped three generations ago can become dead weight, or worse, actively conflict with a newer mechanism, and nothing forces it out unless something is explicitly looking for prunable edits.

RRSI's proposer-selector loop, with an annealed edit budget and novelty search feeding into a critic-and-pruner selection stage, cycling back to update the harness

What the regularization buys you

The authors evaluate across eight benchmarks spanning coding, agentic workspace tasks, and engineering design, with an explicit in-distribution/out-of-distribution split — a harness evolves against one subset and gets scored on both that subset and five benchmarks it never touched during evolution. That OOD split is the part of the experimental design I'd want any team doing this internally to copy, because it's the only way to actually see the failure mode rather than assume it away.

RRSI reaches up to 14.1 points of gain on the benchmarks it evolves against, and — the number that matters — up to 4.7 points of gain on the five OOD benchmarks it never trained on. The comparison point is unregularized evolution, which by the paper's framing produces harnesses whose OOD gains shrink or vanish even as ID scores keep climbing. On top of the generalization result, the regularized harness runs on 30% fewer policy tokens than its unregularized counterpart — a direct consequence of the pruner removing edits that bought marginal accuracy at large token cost, and of the annealed budget keeping later-stage edits lean instead of compounding into larger and larger prompts and tool-call chains.

The operational takeaway

If you're running any kind of automated harness search — prompt evolution, tool-selection tuning, memory-schema search — against a fixed eval set, treat that eval set the way you'd treat a training set for a model: assume overfitting is happening unless you're measuring held-out performance to check. The three levers here are all cheap to add to an existing evolutionary loop: cap and anneal how much a single candidate can change at once, bias the search toward unexplored edits instead of re-scoring near-duplicates of your current best, and add an explicit pruning pass that asks not just "did this help" but "is this still worth its cost." The token-efficiency result is arguably the more durable of the two findings for production use — a harness that's 30% cheaper to run at comparable or better generalized accuracy is a win you keep even if your benchmark suite changes next quarter. Code and a project page are linked from the paper for anyone who wants to look at the proposer and selector implementations directly.

References
  1. 01RRSI: Regularized Recursive Self-Improvement of Agent Harnesses (arXiv:2609.24972)