Scaling doesn't teach a model to catch a falling knife
ReactHuman is a new benchmark that tests whether multimodal LLMs can react in real time to sudden physical hazards, not just answer questions about them — and finds that models mishandle roughly one hazard in three, with no improvement from scale.

I've spent the last few years building agentic AI systems where a wrong output isn't just a bad paragraph — it's a bad action with a consequence downstream. So a new benchmark called ReactHuman, posted this month by Yizhan Li and seven co-authors, caught my attention for a narrow but pointed reason: it doesn't ask a multimodal LLM to describe physics. It asks the model to react to physics, in real time, through a simulated body that pays for being wrong.
A benchmark for reflexes, not reasoning
Most physical-understanding evals for MLLMs fall into one of two buckets. Either they probe intuitive physics passively — question answering over a video of something falling or colliding — or they test deliberate, long-horizon competence like navigation and object rearrangement, where the model has seconds or minutes to plan. Neither measures the layer that sits between those two: the reflex layer, where a household robot's decision core has to turn "I understand this plate is falling" into a committed motor action before the plate hits the floor. ReactHuman is built specifically to isolate that layer. The evaluated MLLM plays the brain of a simulated humanoid that gets ambushed by sudden household hazards — a slipping plate, a falling knife — and has to commit to a reaction with no time to deliberate the way a planning benchmark would allow.
A physics engine, not a human annotator, writes the ground truth
The benchmark spans 17 event families and more than 1,000 bit-for-bit reproducible scenes, with ground truth derived directly from 240 Hz rigid-body simulation rather than human labeling. That matters more than it sounds: there's no annotator judgment call about what "the right reaction" was, because the physics engine already computed exactly where the object goes, how fast, and what intercepts it in time.
The benchmark also includes adversarial objects whose appearance contradicts their physics — a foam anvil, a steel apple. That's a deliberate trap for models that key off visual priors instead of the object's actual motion.

Grading the reaction, not just the output
Every committed plan is physically executed in the simulation, so the decision has an observable, measurable consequence rather than being scored as text. ReactHuman then grades that consequence with a five-metric suite along three axes: whether the reaction was reasonable, whether it was safe, and whether it was physically grounded.

Where seven models actually broke
The authors evaluated seven representative MLLMs against this harness, and the failure modes are specific enough to be useful as a diagnostic rather than just a leaderboard number. Models mishandle roughly one hazard in three. They tend to act from fixed dispositions rather than the scene actually in front of them — meaning the same model reaches for the same reaction across different hazards, as if it were pattern-matching to "a thing is falling, therefore do X" instead of reading the specific trajectory, mass, and timing in front of it. They trust appearance over motion, which is exactly what the foam-anvil/steel-apple objects are designed to expose: a model that reacts to what an object looks like rather than how it's actually moving gets the wrong answer on the adversarial cases even when it would have gotten a normal anvil or apple right. And even when a model picks the correct type of action, it frequently misses the interception point by a meter or more — the right idea, executed against the wrong physics.
The finding that matters: none of this shrinks with scale
The result I keep coming back to is that none of these failure modes shrink with model scale. That's the tell. If the problem were a knowledge gap — models not "knowing" enough physics — bigger models with more training data would chip away at it, the way scale reliably improves benchmarks that measure declarative or reasoning competence. Flat performance across scale is what you'd expect from a different kind of failure: not missing knowledge, but a missing control loop. Reactive safety needs perception and action coupled tightly enough to fire within the time budget of the event itself, using motion cues rather than static appearance. That's closer to what a fast control policy does than to what a language model is trained to do, and it's a capability that "answer more physics questions correctly" doesn't obviously train.
Why this matters if you build embodied or time-critical agents
The practical takeaway, if you're building robotics or any agent that has to make consequential decisions on a clock, is not to assume a stronger foundation model buys you better reflexes. Reactive, safety-critical competence needs to be measured on its own terms — separately from whatever your model scores on static QA or long-horizon planning benchmarks — because ReactHuman's results show those numbers don't transfer. That's true even outside robotics. Any production agentic system that has to act fast under uncertainty, with real consequences attached to being wrong, deserves the same kind of test: not "does the model know the right answer," but "does it act correctly when the scene, not the prompt, is what's changing." ReactHuman is a useful template for building that test, and its central result — that today's MLLMs fail this in ways scale doesn't fix — is a reason to build it before you need it.