Writing
September 19, 2026 · 6 min read

SoL-Pi: an automated search over agent-harness design cuts token cost by a third — with real caveats

A new NVIDIA/NTU/MIT paper runs an automated search over agent-harness design choices rather than the model itself, surfacing four mechanisms that cut coding-agent token traffic up to 49% and cost by about a third — at a real, if small, accuracy cost the abstract doesn't emphasize.

agent-harnesstoken-efficiencyautomated-researchllm-agentsproduction-ai

What actually shipped

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness comes from a team at NVIDIA, NTU, and MIT — Song Han is among the fourteen authors — and its premise is narrower and more useful than the "recursive self-improvement" framing in the title suggests. They didn't get an agent to improve a model. They got an agent to audit another agent's harness — the code that decides how tool calls get batched, how context gets trimmed, and how log output gets summarized — and ran a large automated search to find which harness changes actually save tokens without costing accuracy.

The headline result: on the public 51-task slice of EdgeBench, the resulting harness, SoL-Pi, cuts recorded token traffic 44.7–49.0% and API cost by about a third relative to Pi, an existing open coding-agent harness, while retaining 93.7–94.3% of Pi's task score. That "retaining" is doing more work than the abstract lets on, and I'll get to why.

Four mechanisms, not one clever trick

SoL-Pi is four extensions bolted onto Pi, each targeting a different place tokens get wasted in a long-running coding agent:

  • Action Fusion — Pi normally edits a file, then issues a separate command to test or build it: two model round-trips. Action Fusion merges the edit and its follow-up command into one tool call and returns both outcomes in a single observation, cutting three API calls to two for that step.
  • Online Context Compact — instead of compacting on a fixed schedule, the harness checks at every completed plan step whether the projected token savings exceed the cost of rewriting the prompt cache, and only compacts when that gate passes.
  • ObservationPack — large tool outputs (over 10 KiB) go out in full for the first two follow-up requests, then get replaced with a stable handle plus a 1 KB excerpt; the agent can still pull the exact original on demand.
  • Evidence-Preserving Reducer — build and test logs over 4 KiB get routed through a cheaper model that extracts verified evidence into a compact receipt, checked by a deterministic verifier, with a fallback to the raw log if verification fails.

Auto-research loop narrowing 152 proposed harness changes through 535 search environments down to four retained mechanisms — Action Fusion, Online Context Compact, ObservationPack, and Evidence-Preserving Reducer — combined into the SoL-Pi harness for a 44.7–49.0% token cut and about a 33% cost cut.

None of these four ideas is conceptually new on its own — the paper's own related-work section cites half a dozen prior systems doing context folding, trajectory reduction, or log compression. What's distinct is that none of the four was hand-designed. They're what survived an automated search over roughly 150 candidate harness changes.

How the search worked

This is where the RSI framing earns some of its keep. A research agent inspects execution traces from a separate agent running the base harness, proposes a harness change, implements it, and tests it in one of 535 executable environments — 495 derived from real GitHub issue/pull-request pairs (fix hidden, regression test required to fail before the patch and pass after) and 40 synthetic tasks built around executable verifiers. Every candidate has to clear two fixed gates: it can't degrade capability beyond a preset tolerance, and it has to improve at least one declared efficiency metric. Across roughly 3,000 runs and more than 60,000 agent-environment interactions, four mechanisms survived.

The numbers, and what they cost you

On EdgeBench with GPT-5.6 Sol, the efficiency-tuned SoL-Pi configuration runs 1.10B total tokens against Pi's 2.15B — a 49.0% cut — at $894 versus $1,339 in API cost, while scoring 42.0 against Pi's 44.8 (93.7% of Pi's average). Transferred untouched to Opus 5, a backend the harness was never tuned against, it holds 94.3% of Pi's score while cutting tokens 44.7% and cost 33.5%. That cross-model transfer is the paper's strongest evidence the discovered mechanisms generalize rather than overfitting to one model's tool-calling habits.

But the paper is candid about where the pattern breaks. On the held-out Opus 5 backend, the four mechanisms "trigger less often and less intensively" than they did on the model they were mined from — a visible sign the search partly learned GPT-5.6 Sol's specific action patterns, not just general harness inefficiency. And on Terminal-Bench 4, SoL-Pi solves 15 of 63 CPU-only tasks against 18 for both Codex and Pi — fewer tasks outright, at a lower cost per task solved ($14.07 vs. $15.91 for Pi). On IMO 2026, it passes 3 of 6 problems, tying Pi and trailing Codex's 5, again at the lowest cost per passed problem. "Comparable quality," the phrase the abstract leans on, holds on average and in aggregate score; it does not hold task by task.

The methodologically important part is what happens before those numbers get reported: EdgeBench is frozen out of the search entirely. Candidates are selected on separate development environments, then evaluated once, cold, on held-out tasks that never feed back into the loop. That separation is a direct response to a problem the paper cites by name — other researchers (Wang et al., 2026) found that harnesses evolved against an eval set can overfit it and show only marginal gains on genuinely unseen tasks. SoL-Pi's authors clearly designed around that failure mode instead of ignoring it, which is more methodological discipline than most harness-optimization papers show.

My take

The most credible sentence in this paper isn't in the abstract — it's in the limitations section. The authors call "recursive efficient improvement," the idea that a cheaper harness could fund the search for its own successor, "a long-term research vision rather than a compounding effect demonstrated by the present study." That's an unusually direct admission from a paper whose title promises "recursively scaling" loops, and it tells you the real contribution is narrower than the framing: a disciplined, held-out-validated search that found four specific, reusable harness-level fixes, not evidence of anything actually compounding yet.

Those four fixes are worth stealing even if you never run the search. If you operate a production coding agent, the transferable lesson is a checklist, not a system to install: are you issuing a separate API call to verify every edit you just made? Are you compacting context on a timer instead of a cost/benefit gate? Are you replaying full tool output every turn instead of archiving it behind a handle? Are your build logs going into the model raw instead of through a cheap, verified extractor? Those four questions are model-agnostic, and this paper is decent evidence that answering them well is worth roughly a third of your token bill — as long as you're willing to give up a few points of task score for it, which is the trade nobody puts in the headline.

References
  1. 01SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness (arXiv:2609.20519)
  2. 02EdgeBench
  3. 03Pi: Coding Agent Toolkit (Earendil Works)
  4. 04Terminal-Bench 4.0
  5. 05Rethinking the Evaluation of Harness Evolution for Agents (Wang et al., arXiv:2607.12227)