DeepSeek-V4.1-Flash targets the bottleneck long-context agents actually hit: KV cache memory
DeepSeek's new V4.1-Flash cuts KV cache footprint to 890 bytes per token — a quarter of its predecessor in HBM, an eighth on SSD/host memory — through cross-layer cache reuse, FP4 quantization, and an asymmetric encoder/decoder architecture built for input-heavy agentic workloads.

Most of the public conversation about frontier models is still about benchmark scores. DeepSeek's latest paper is about something that matters more once you're actually running agents in production: how much memory it costs to keep a model's attention state alive.
DeepSeek-V4.1-Flash is a 552B-parameter multimodal Mixture-of-Experts model supporting contexts up to one million tokens. The headline number in the paper isn't a benchmark score — it's 890 bytes. That's the KV cache footprint per token the model holds in HBM, roughly a quarter of what its predecessor, DeepSeek-V4-Flash, required. On SSD or host memory, where long-lived agent sessions get paged out between turns, the persistent footprint drops to roughly an eighth. The model is already on Hugging Face, where it's pulled in nearly 430K downloads and over 3,000 likes in about a week, plus a fast-growing ecosystem of community GGUF and quantized forks — the kind of adoption curve that tells you people are actually deploying it, not just reading the abstract.
The bottleneck agentic AI actually hits
The paper's framing is worth sitting with: "the widespread adoption of long-horizon agents has made model workloads increasingly input-heavy." That matches what I see running agentic systems day to day. An agent's context isn't a single prompt — it's a conversation history plus every tool call result, every retrieved document, every intermediate reasoning step, accumulated turn over turn. The generation itself (the decode side) is often short. The input (the prefill side) is where the token count explodes.
Every one of those input tokens has to be represented in the KV cache — the per-layer, per-head key and value tensors the model needs to attend back over everything it's already seen. That cache scales linearly with context length, and at a million tokens it becomes the thing that actually limits deployment: how much HBM a GPU has left over for concurrent sessions, how much has to be evicted to slower storage, and how much bandwidth it costs to page it back in when a session resumes. Compute has gotten a lot of optimization attention over the last two years. Memory, for agentic workloads specifically, is now the harder constraint.
An architecture built around the asymmetry
DeepSeek-V4.1-Flash's first move is architectural. It uses what the paper calls a Causal Encoder-Decoder (CED) design that activates only 8B parameters per token during prefill but 16B during decode. That split is a direct response to the input-heavy shape of agentic traffic: prefill has to chew through potentially hundreds of thousands of input tokens, so keeping the active parameter count low there matters enormously in aggregate. Decode generates comparatively few tokens per turn, so spending more active capacity per token there is nearly free in total cost, while buying back quality. It's a deliberate mismatch between the two phases of inference, sized to where the tokens actually are.
Three techniques stacked on the KV cache itself
The architecture handles compute asymmetry; the rest of the paper is about shrinking what has to be stored. Three techniques stack together:
CSA2 (Compressed Sparse Attention 2) introduces cross-layer KV cache reuse — instead of every transformer layer holding its own independent key/value tensors, layers share compressed representations. That multiplies savings across the depth of the network rather than compressing each layer in isolation.
FP4 KV caching stores the cached keys and values in 4-bit floating point instead of the 16-bit (or even 8-bit) precision typical of production serving today. That's a straightforward multiplier on top of the cross-layer savings — fewer bits per stored element, on top of fewer elements needing to be stored at all.
SWA Bounded Replay is the one aimed specifically at the persistent cache — the copy that sits on SSD or host memory for sessions that aren't actively being served. The name points to the mechanism: sliding-window attention layers only need a bounded, recent window of context to attend over, so instead of persisting the full history for those layers indefinitely, the system can cap what's stored and recompute ("replay") anything outside the window if it's ever needed again. That's what gets the persistent footprint down to roughly an eighth of V4-Flash, a bigger win than the HBM-side quarter, which makes sense — persistent storage is exactly where unbounded conversation history would otherwise accumulate without limit.

What 890 bytes a token buys you
Run the arithmetic forward and the number stops being abstract. At the model's full one-million-token context, 890 bytes/token works out to roughly 890MB of HBM for a single sequence's KV cache — down from an implied ~3.5GB for V4-Flash at the same length, using the paper's own 4x ratio. That difference isn't cosmetic. It's the difference between a GPU that can hold a handful of million-token agent sessions in memory at once and one that can hold four times as many, or between a session that stays resident in HBM and one that gets evicted and has to be paged back in from SSD mid-conversation, which is exactly the latency spike users notice.
What's notable is that DeepSeek reports this compression comes without a quality tradeoff — the paper states the model "delivers substantially better performance than the baseline" despite the much smaller cache. If that holds up under independent evaluation, it means the compression isn't a lossy shortcut being traded against capability; it's closer to removing redundancy that didn't need to be there in the first place.
My take
This is a systems paper wearing a model-release paper's clothes, and that's exactly why it's worth paying attention to. DeepSeek has a track record of treating KV cache and inference efficiency as first-class problems rather than afterthoughts, and this release continues that pattern rather than chasing another benchmark table. For anyone running long-context agents in production, the constraint you hit first usually isn't "is the model smart enough" — it's "how many of these sessions can I actually keep in memory at once, and what does it cost when one gets evicted." A model that's a quarter the KV footprint in HBM and an eighth on persistent storage, at comparable or better quality, changes that math directly: more concurrent long-context sessions per GPU, less eviction churn, and a lower floor on what long-horizon agentic deployment costs. That's a more durable kind of progress than another point of benchmark accuracy, and it's the kind of contribution that tends to get adopted quietly and broadly rather than loudly — which the download numbers already suggest is happening.