Writing
October 1, 2026 · 6 min read

Your eval pipeline has an undocumented input: today's date

A new paper shows that the current date silently injected into LLM system prompts shifts benchmark scores by up to 14% on math reasoning and reshuffles leaderboard rankings — a reproducibility hole most eval harnesses never control for.

llm-evaluationreproducibilitybenchmarkingsystem-promptsresearch

I've spent enough time staring at eval dashboards to know the drill: a score moves, you check the diff, you check the seed, you check batch size and numerical precision, and if none of those explain it you shrug and chalk it up to noise. A paper that landed this week gives that shrug a name, and it's not a flattering one.

Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation, by Mario Sanz-Guerrero, Minh Duc Bui, Manuel Mager, and Katharina von der Wense (accepted to AACL 2026), identifies a confound that's been sitting in plain sight: most providers inject the current date into the system prompt automatically, on every call, whether or not the user asks for it. You don't see it, you can't turn it off, and it changes every single day. The paper asks the obvious follow-up question nobody had actually tested: does that silently-injected date change what the model outputs? Across 9 recent LLMs and 6 datasets spanning multiple-choice QA, math reasoning, code generation, and machine translation, the answer is unambiguously yes.

What they actually measured

The setup is clean. Take a fixed model, a fixed prompt, a fixed question — the only thing that varies is the date string the system prompt claims is "today." Run the same evaluation across different injected dates and watch the score move. No code change, no retraining, no different checkpoint. Just a different date in a part of the prompt most of us never look at because we didn't write it.

The deltas are not rounding error. The authors report swings of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation, purely as a function of which date was injected. Fourteen points on math reasoning is bigger than the gap that separates a lot of model generations from each other on public leaderboards. And because different models react to dates differently, the effect doesn't just add noise — it changes relative rankings. A model that beats another on Tuesday's date string can lose to it on Friday's, with nothing else different about the comparison.

The part that should worry anyone running evals for a living: the paper checks whether the usual mitigations help, and they don't. Chain-of-thought prompting doesn't damp the sensitivity — it amplifies it. Few-shot prompting doesn't fix it either. The techniques we reach for to make model behavior more stable and more reasoned turn out to make this particular instability worse, not better, because the model has more room to let the date leak into its reasoning trace before it ever produces an answer.

Side-by-side diagram showing the same model and question producing 61% math accuracy with one injected date and 47% with another, a 14-point swing with no code change, plus a bar comparison showing the date effect exceeds batch-size and precision noise

Why this is worse than the noise sources we already track

Reproducibility problems in LLM evaluation aren't new — prior work has already documented that outputs drift with hardware, batch size, and numerical precision. Those are annoying, but they're also the kind of thing a careful team can control for: pin the hardware, fix the batch size, log the precision. The paper's headline finding is that the date effect is larger than those known sources of non-determinism, and it is categorically harder to control because it isn't exposed as a parameter at all. You can't pin a variable you don't know exists. It's not in your eval config, it's not in your model card, and unless you're logging full raw requests and diffing system prompts by hand, it's invisible until a score moves for no reason you can find.

There's also a mundane but important detail buried in the mechanism: the model plausibly uses the date as a cue about its own training cutoff, recency, or what's likely to be "current" in the question's context, and that cue bleeds into answers on tasks that have nothing to do with dates — a math word problem, a translation, a multiple-choice question. The date isn't content. It's metadata about the request. But the model doesn't reliably treat it that way.

What I'd actually do about this

If you maintain an eval harness or a leaderboard, the practical takeaway is to stop treating the system prompt as a constant. Log the full system prompt — not just your own template, the final one the provider actually sends — for every eval run, and record the date string if one is present. If you're comparing two model runs, run them with the same injected date, or better, sweep a range of dates and report a distribution instead of a point estimate. A single-date benchmark score is now a biased sample of one.

If you're building an agent that calls these APIs in production, the implication is narrower but still real: don't assume determinism across days for prompts that are sensitive to "current" framing, and don't be surprised if a regression you're chasing down turns out to correlate with the calendar rather than your last deploy. It's a cheap first check, and after reading this paper, exactly the kind of thing I will be looking at before I blame my own code.

The broader point is one I keep relearning in this field: the inputs you didn't choose are still inputs. A hidden date is a small thing to add to a system prompt, and it moves scores more than batch size does. Worth checking what else providers are quietly appending to those prompts that we've been treating as fixed.

References
  1. 01Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation (arXiv:2609.36931)