Writing
September 13, 2026 · 6 min read

How T1 closes the training-inference gap that breaks long-horizon agentic RL

A new 122B-parameter MoE terminal agent jumps from 43.8% to 64.0% on Terminal-Bench 2.1 by fixing a subtle mismatch between what gets sampled during rollout and what gets trained on afterward. The mechanism, not the score, is the useful part.

reinforcement-learningagentic-aimixture-of-expertsterminal-agentsllm-training

Most agent benchmarks fall apart within a few dozen tool calls: the model loses track of what it already tried, the sandbox drifts from its internal picture of it, and a mistake from turn five doesn't get corrected until turn twenty, if it ever does. T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks is aimed squarely at that failure mode. T1 is a 122B-parameter mixture-of-experts (MoE) model, post-trained with reinforcement learning while operating a real shell in a cloud sandbox, on tasks that can run past 300 tool-call turns before the task's own verifier grades the result. The headline is a jump from 43.8% to 64.0% resolved on Terminal-Bench 2.1, with 27.9% on the harder Long-Horizon Terminal Bench, ahead of GPT-5.4 and GLM-5.1 on that split.

The score is not the interesting part. The interesting part is buried in the infrastructure section, where the authors report cutting the training-to-inference log-probability gap from 0.021 to 0.013 and reaching exact zero token drift in the region that determines the loss. That's a narrow, almost obscure-sounding metric, but it's the thing that decides whether long-horizon RL on a MoE model converges to something useful or quietly poisons itself over the course of training. I want to spend most of this post on why that gap exists and how the paper closes it, because the fix generalizes well past this one benchmark.

Why 300 turns breaks ordinary on-policy RL

On-policy RL for language models has a structural assumption baked in: the policy that generates a trajectory (the sampler) and the policy whose gradients you compute afterward (the trainer) are supposed to be the same function. In practice they never quite are. The sampler usually runs on an inference-optimized stack — different batching, different kernels, sometimes different numerical precision — while the trainer recomputes log-probabilities over the same sequence using its own forward pass. For a single short generation, the resulting numerical mismatch is small enough to ignore. Stretch the trajectory to 300+ tool-call turns in a real shell, and small per-token discrepancies compound. The gradient you compute is no longer quite the gradient of the policy that actually produced the reward.

Mixture-of-experts models add a second, separate axis to this problem. A MoE layer routes each token to a subset of experts, and that routing decision is itself a function of the input and the exact numerics of the forward pass. If the trainer recomputes routing independently of the sampler — which is the default behavior in most RL training setups — it can send a token through a different set of experts than the ones that actually generated it during rollout. At that point the trainer isn't just computing gradients with slightly noisy log-probabilities; it's computing gradients for a structurally different computation than the one that earned the reward. Over a handful of turns this is a rounding error. Over 300+ turns of terminal commands, retries, and verifier checks, it becomes the dominant source of noise in the training signal.

TITO: train on the tokens you actually sampled

The first fix, which the paper calls TITO, is to stop letting the trainer regenerate its own version of the trajectory. Instead, T1 trains on the exact sampled token identifiers produced during rollout, with drift repair applied at turn boundaries — the points in a long shell session where accumulated context, retokenization, or environment state could otherwise let the trainer's view of the sequence diverge from what was actually executed. It's a conceptually simple constraint: don't let the two halves of the RL loop disagree about what happened.

R3: replay the router instead of recomputing it

The second fix, rollout routing replay (R3), does the equivalent for the MoE router. During rollout, T1 records the sampler's per-token expert choices at every MoE layer. During training, instead of letting the trainer recompute those routing decisions from scratch, it replays the recorded choices. The trainer's forward pass then matches the sampler's forward pass not just in the tokens it sees but in the literal compute path those tokens took through the model.

Diagram comparing standard RL training, where the trainer recomputes tokens and MoE routing and drift accumulates over 300+ turns, against T1's TITO and R3 approach, which replays the sampler's exact tokens and routing and holds the log-probability gap near zero

Together, TITO and R3 are what take the training-to-inference log-probability gap from 0.021 down to 0.013, with the paper reporting exactly aligned zero token drift in the loss region. Neither piece is exotic on its own — recording and replaying state is an old idea — but applying it consistently to both token identity and expert routing, across a training loop that spans hundreds of tool calls, is what makes long-horizon RL on a MoE agent stay stable instead of drifting into a policy that no longer matches what was actually rewarded.

An out-of-distribution training corpus, on purpose

The paper also takes a deliberate step to guard against a different kind of failure: benchmark overfitting. The RL training corpus uses isolated seeds and synthesized tasks that are disjoint from Terminal-Bench 2.1 itself. Combined with a dense process reward — trajectories are scored by the absolute number of passing verifiers rather than a single pass/fail signal — and an aggressively warm-started actor-critic setup for training stability, this is meant to ensure the reported gains reflect transferable capability rather than the model having implicitly memorized the shape of the eval it's graded on. Given how easy it is for agent benchmarks to be gamed by training too close to the test distribution, this is worth taking seriously as a methodological choice, not just a footnote.

My take

What I find genuinely useful about this paper is that it treats the sampler/trainer mismatch as a first-class engineering problem rather than noise to average away. That mismatch is not unique to terminal agents — anyone running on-policy RL on a MoE model over multi-turn tool use is exposed to the same drift, whether the domain is coding, browsing, or scientific workflows. The fact that it takes 300+ turns to make the problem visible in this paper's numbers doesn't mean shorter agentic RL runs are immune; it means shorter runs are less likely to have accumulated enough drift to show up in the metrics yet. If you're building or fine-tuning agents against long tool-use trajectories on a MoE base model, the token-identity and routing-replay discipline described here is the kind of infrastructure detail that decides whether your RL run converges or slowly degrades in ways that are hard to diagnose from the loss curve alone. It's not a glamorous contribution, but it's the kind that tends to matter more, in practice, than another point of benchmark score.

References
  1. 01T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks