Register tokens: giving diffusion language models a bounded memory for reasoning
A new paper trains diffusion language models to carry reasoning across generation chunks in a fixed-size set of register tokens instead of the growing text context, beating a discrete-text-carry baseline by up to 19.5 points on code and 8.5 on math.

I've spent enough time optimizing context windows in production systems to have a healthy suspicion of any architecture whose memory cost scales with how much it's already said. A new paper, Register Tokens for Bounded-State Reasoning in Diffusion Language Models (Ge, Singh, Zhuang, Liu, Gao, and Sala), attacks exactly that problem for masked diffusion language models, and the fix is small enough to describe in one sentence: replace the growing text context with a fixed number of hidden-state tokens that get carried forward and nothing else.
The problem is specific to how diffusion LMs generate
Masked diffusion language models (dLLMs) like LLaDA and Dream don't generate left-to-right the way autoregressive transformers do. They start from a block of masked tokens and iteratively denoise it with bidirectional attention until the block resolves into text. That's a genuinely different generation mechanism, and it's part of why dLLMs are attracting attention as an alternative to the standard decoder-only recipe.
The catch shows up when you need output longer than one chunk. If a proof or a program spans multiple denoising rounds, the model has to keep reasoning coherent across chunks. The default way to do that is the same trick autoregressive models use: keep the previously generated text in context and condition the next chunk on it. It works, but the context grows with every chunk, and everything the model produced earlier — including tokens that were only scaffolding for the final answer — stays in the window forever.
The fix: a small set of registers instead of a memory that grows
The paper's proposal is to add a fixed number of dedicated positions — register tokens — whose continuous hidden states are trained to hold reasoning progress. The model decodes a chunk of text as usual, but then that text is cleared from context while the register states are preserved. The next chunk is decoded from just the original prompt plus those carried register values, not from the accumulated text history.
That's the whole mechanism, and the interesting part is training, not architecture. The dLLMs are post-trained specifically to decode a chunk, discard the visible text, and continue reasoning from the prompt and the carried registers alone. The registers aren't a cache of past tokens; they're a compressed, continuous encoding the model has learned to write to and read from, closer in spirit to a fixed-size working memory than to a context window.

What the comparison actually measures
The paper's main baseline is what I'd call the obvious approach: discrete text carry, where the previous chunk's generated text is kept verbatim in context for the next chunk. Both methods start from the same LLaDA and Dream checkpoints, so the comparison isolates the representation used to carry state, not the base model.
On that comparison, registers outperform discrete-text carry on every benchmark tested, with gains reported up to 8.5 points on math and 19.5 points on code. The code result is the one worth sitting with. The authors note registers are especially effective for bounded code generation, where a correct program typically has to span several chunks and a single broken thread — a variable defined in chunk one and misused in chunk three — breaks the whole thing. A fixed-size state that's been trained specifically to preserve reasoning continuity, rather than a growing pile of raw text the model has to re-parse, seems to hold that thread better.
It's also worth noting the registers aren't just a distillation target frozen after pretraining: the paper reports they can be further refined with reinforcement learning on long-horizon reasoning tasks, which suggests the carried state is treated as a first-class object in the training loop, not a side effect of chunking.
Why this matters beyond diffusion LMs
The framing that stuck with me is the question the authors ask directly: can a model continue reasoning after the text that produced that reasoning is gone? For autoregressive transformers, that's mostly a non-question — attention over the KV cache is the mechanism, so the state is the text, and there's no clean way to separate them without approximation (compression, summarization, or retrieval, all of which are lossy in different ways). The bidirectional, chunked structure of dLLMs gives you a genuine seam to insert a bounded carried state, and this paper's result is a reasonably clean demonstration that the seam is useful, not just possible.
Practically, this is a data point for anyone in the position I'm often in: running generation over long, multi-step outputs where context cost and coherence trade off against each other. Discrete text carry is what happens by default when you extend a chunked generation process — you just don't throw anything away. This paper's evidence is that a purpose-trained fixed-size state can substitute for that growing context without a coherence penalty, and in the code setting, with a fairly large one in the other direction. It's a narrow result — one paper, two base models, a specific benchmark suite — but it's the kind of narrow result that tends to generalize once someone tries it on a third model family.
I'd want to see the register mechanism tested at longer horizons than the paper's chunk counts before calling it a settled technique, and I'd want to know how the fixed register size interacts with problem difficulty — does an eight-register state saturate on a proof that needs to track more than eight live facts? Those are the questions I'd chase first if I were extending this work, and they're the reason I'll be watching whatever follow-up comes out of this group.