Writing
September 25, 2026 · 6 min read

MindTopo shows vision-language models can see structure but can't hold onto it

Microsoft Research's new MindTopo benchmark finds that multimodal models recognize topological relationships like connectivity and enclosure in a static image, but lose track of them once they have to act on a scene over multiple steps.

vision-language-modelsspatial-reasoningbenchmarksroboticsmicrosoft-researchagentic-ai

Microsoft Research just published MindTopo, a benchmark that tests something most spatial-reasoning evaluations skip entirely: topology. Not distance, not bounding boxes, not "is the cup to the left of the plate." Topology is the layer underneath that — whether two regions stay connected, whether a boundary keeps something in or out, whether a loop of rope is actually knotted or just tangled to look that way. It's a property that survives bending and stretching, which is exactly why it matters for robots and agents operating in real, deformable environments rather than static photographs.

Why topology, specifically

The MindTopo team — researchers from Northwestern, Stanford, and Microsoft — grounds the benchmark in Piaget's classification of topological cognition, organizing tasks into five categories: continuity (does a path stay unbroken), separation (are elements one structure or distinct parts), order (how elements are arranged along a transformation), enclosure (does a boundary create an inside and an outside), and knots (is something truly knotted or merely tangled). Each category gets tested two ways: a reasoning task, where the model looks at one or more rendered scenes and answers a question about their structure, and a planning task, where the model has to act — rotating pipes, drawing a separating path, untangling ropes — inside a simulator that enforces legal moves, so there's no cheating by passing one strand through another.

That reasoning/planning split is the interesting design choice. It lets the researchers separate two failure modes that look identical from the outside: a model that fails because the scene is visually cluttered, and a model that fails because it can't hold onto a relationship once things start moving.

MindTopo's five topological test categories and the gap between static reasoning and sequential planning performance

What actually broke

Across a wide set of proprietary and open-weight models, the pattern was consistent: strong-ish performance on static recognition, a sharp drop on interactive planning, and both bands sitting well below human baselines. The error analysis is the useful part. Static failures were mostly perceptual — a model missed a wall, an opening, a crossing point. Planning failures showed up after the model had already understood the scene correctly. It would take a locally reasonable move, then lose track of the consequences a few steps later, or propose an action the environment's own physics wouldn't allow. In other words, the bottleneck isn't seeing the topology. It's carrying it forward.

The team also tested whether generative image and video models could substitute for explicit state-tracking — essentially, can the model "imagine" the next frame accurately enough to reason from it. Image generation helped a little when the relevant relationship was visible in a single frame, but broke down across longer sequences of moves. Video rollouts frequently altered the topology outright or violated task dynamics mid-generation. The tools were only as useful as their ability to preserve structural constraints over time — which, notably, is the exact same failure they were being asked to compensate for.

Why this matters beyond the benchmark

I spend most of my working time on agentic systems, and this result lands squarely on a problem I've hit in practice: an agent that's individually correct at each step can still be globally wrong, because nothing in the architecture is tracking what has to remain true across steps. MindTopo gives that intuition a controlled, quantifiable form. A model that can label "the sheep is inside the fence" from a photo has learned a pattern-matching shortcut, not a representation of enclosure that survives the fence moving.

The paper's own framing of the fix is worth taking seriously: closing this gap likely requires either models that carry an explicit topological state alongside their token stream, or world models whose predictions preserve topology by construction rather than by luck. Neither is a prompting trick. Both are architecture questions. For anyone building robotics stacks, spatial assistants, or long-horizon agents on top of today's VLMs, MindTopo is a reasonable diagnostic to run before assuming perception parity means planning competence — they're measurably not the same thing.

References
  1. 01MindTopo reveals VLMs' spatial reasoning abilities — Microsoft Research Blog