Curate memory when the task is known, not when the task ends
A new paper argues that agent memory systems have been curating at the wrong moment — distilling experience into fixed summaries right after a task ends, instead of waiting until a new task arrives and shaping memory around it. The fix trains cleanly on immediate task success instead of a delayed reward signal.

Every agent memory system I've shipped or reviewed makes the same architectural bet: decide what's worth remembering right after a task finishes, while the outcome is fresh, and store that distillation — a reflection, a workflow, a reusable skill — for retrieval later. Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents (Zhou, Li, Liu, Yavuz, and Joty) argues that bet is backwards, and the argument is tight enough that I think it changes how I'd design the next memory layer I build.
The problem with curating at write time
Call the standard approach write-time curation. A task completes, an agent (or a separate curator model) looks at the trajectory and decides what to keep: a reflection on what went wrong, a generalized workflow, a distilled strategy. That artifact gets embedded and stored. Later, a new task comes in, the system retrieves artifacts by similarity, and the agent conditions on whatever got kept.
The paper names two structural problems with this, and both match what I've run into in production agent systems:
You're guessing at the future query. Curation happens before you know what the next task will actually need. So the curator has to produce something generic enough to serve many hypothetical downstream tasks — which means it's optimized for nothing in particular. Information that would have been exactly right for a specific future task gets thrown away because it didn't look generally useful at write time.
Credit assignment is long-horizon and indirect. If you want to learn a good write-time curator — rather than hand-write heuristics — you need a training signal. But the value of a storage decision only becomes visible when a similar query shows up, which might be dozens of tasks later. That's a classic delayed-reward problem, and it's exactly the kind of signal that's hard to train against reliably.
The fix: defer curation until the task is known
JitMem's answer is to stop compressing at write time at all. It retains raw trajectories — the full, uncompressed experience — and defers curation to read time, once the new task is actually in front of the system. Given the retrieved raw traces and the current task, a curator model synthesizes a compact, task-adaptive payload tailored to that specific need, on the spot.
This reordering is what makes the training problem tractable. Because the payload is synthesized for, and consumed by, the same task it's built for, the curator can be trained directly from immediate task success — did the agent succeed with this curated payload, yes or no — instead of a distal, multi-hop credit-assignment signal. The feedback loop collapses from "maybe helpful several tasks from now" to "helpful on this task, right now," which is a signal you can actually optimize against with standard RL.

It's worth being precise about what moved. Retrieval — pulling candidate traces by similarity — still happens the same way it always has. What changed is what happens after retrieval and before the agent acts: instead of handing over a pre-baked artifact, the system runs a synthesis step conditioned jointly on the raw traces and the live task, and only that synthesis step is what's learned.
What the results say about where the gain comes from
The authors evaluate on three established agent benchmarks — ALFWorld, WebShop, and τ²-bench — comparing against no-memory agents and both heuristic and learned write-time memory baselines. JitMem beats the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points on the three benchmarks respectively. Those aren't marginal gains; on ALFWorld and WebShop they're the kind of jump that changes whether a system is usable.
The more interesting result, though, is a decomposition. An untrained curator — read-time synthesis with no RL applied at all — is already competitive with or beats the write-time baselines. Training the curator further compounds the improvement, but it isn't where most of the gain originates. That tells you the win is mostly architectural: just moving curation to the moment the task is known, so the summary is built for a concrete need rather than a hypothetical one, does most of the work. Learning refines it further, but the shift in when you curate is the load-bearing decision, not the shift in how you optimize the curator.
Why this matters if you're building long-running agents
If you're running an agent that accumulates experience over many sessions — a coding agent, a support agent, a research assistant with a growing memory store — this reframes a decision most teams make implicitly and never revisit: compress on write, retrieve on read. That default is convenient (you pay the curation cost once, storage is cheap and structured, retrieval is fast) but it locks in the two failure modes above by construction.
The practical implication isn't necessarily "go implement JitMem verbatim." It's that the design space is bigger than most memory stacks assume. Storing raw or lightly-processed trajectories and pushing synthesis to read time costs you retrieval-time latency and compute you didn't pay before, but it buys you a task-specific artifact and — crucially — a training signal you can actually shape, since success is observable on the same task the memory was built for. For any agent where the downstream task distribution is broad and unpredictable, that trade looks increasingly like the right one, and I'd expect more memory architectures to move this direction over the next year.