Entities as anchors: giving agentic search a map of the corpus
A new paper, Follow the Entities: A Corpus Map for Agentic Search, builds an offline graph of recurring entities that links documents together so a search agent stops rediscovering the same relationships on every query — and uses fewer tokens doing it.

Most of the multi-hop retrieval problems I run into in production don't look like textbook QA. They look like this: a project's approval lives in a kickoff doc, its requirements sit in a spec that was revised twice, and its current status is buried in a status update from three weeks ago. None of those documents know about each other. An agentic search loop has to discover that they're related by reading, guessing, and re-querying — for every single question that touches that project, because that relationship was never written down anywhere the agent can look it up.
That's the exact gap a new paper addresses: Follow the Entities: A Corpus Map for Agentic Search, by Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, and Andrew Joohun Nam (submitted September 29, 2026). The core move is simple to state and, I think, underused in how we build retrieval systems: stop treating the corpus as a flat pile of files and give the agent a navigation layer built around the entities that recur across it.
The problem with searching a flat corpus
The dominant pattern for agentic search over large document collections is iterative: the agent issues a query, reads what comes back, decides what it's still missing, and issues another query. This beats naive top-k retrieval because it lets the agent chase evidence across multiple rounds instead of betting everything on one ranked list. But the paper points out the structural cost of doing this over a flat collection — when a document gives no indication of how it relates to others, the agent has to rediscover those relationships from scratch every time. Two failure modes follow directly from that: the agent misses complementary evidence sitting in a document it never thought to search for, and it burns tokens re-deriving connections that were true yesterday and will be true tomorrow, because nothing was cached across queries.
That second point is the one that stands out to me operationally. Every query starting cold on relationship-discovery isn't just inefficient, it means the work an agent does to relate document A to document B has zero shelf life. The same rediscovery cost gets paid again on the next query, and the one after that.
CorpusMap: entities as the shared index
The paper's answer, CorpusMap, is built offline, before any query is asked. It resolves entity mentions across the whole corpus — the same project, person, or system referenced under slightly different names in different documents — and for each recurring entity builds what they call an Entity Page: an aggregated summary of what's known about that entity, plus links to every document that mentions it. Do this for enough entities and you get a graph where entities and documents are nodes, and an agent can traverse edges instead of re-searching blindly.
Concretely, this turns "which documents relate to Project X" from a query-time search problem into a graph lookup: land on Project X's Entity Page and the connected documents — the approval, the spec revisions, the status update — are already linked, regardless of what vocabulary each one used to refer to the project. The relationship was resolved once, offline, and is now available to every future query for free.

That offline/online split is the whole trick. Entity resolution is exactly the kind of work you want to do once and amortize — it doesn't depend on any particular question a user will eventually ask, only on what's actually in the corpus. Pushing it out of the query loop is what turns a per-query tax into a one-time build cost.
What the results say
The authors test CorpusMap across 7 different models and 3 benchmark datasets, comparing it against raw-corpus agentic search and against 4 alternative navigation layers. The reported outcome is a clean win on the axes that matter for this kind of system: better evidence discovery, better answer quality, and fewer tokens on average than searching the raw corpus — and CorpusMap also outperforms the other navigation-layer baselines they tried. The paper's framing of that last result is worth sitting with: it suggests entities specifically, not just structure in general, are what make good anchors for navigating a large document collection. A generic graph over the corpus isn't the same as a graph organized around the things that actually recur and matter — people, projects, systems, decisions.
Why this matters if you're building agentic search
I'd separate the contribution into two ideas, because they're useful independently. First: hard-won relational knowledge about a corpus (this document relates to that one, through this entity) should be treated as an artifact you build and store, not a byproduct you regenerate on every query. Second: entities are a good organizing key for that artifact specifically because they're identifiable from the documents themselves — you don't need an external ontology or manual tagging, just resolution across mentions.
For anyone running retrieval-augmented agents over an internal knowledge base — tickets, specs, meeting notes, status docs — the practical implication is that entity resolution isn't just a data-cleaning nicety. Done well and exposed as a navigable structure, it's an evidence-discovery mechanism that pays for itself the moment you have more than a handful of queries touching the same entities. The paper is a strong argument for building that layer offline rather than leaving your agent to reconstruct it, query by query, forever.