A paper that answers questions: Paper2Agent turns research output into an MCP server
A new Nature paper, Paper2Agent, converts a research paper and its codebase into an MCP server that any chat agent can query — reproducing figures, running the method on new data, and citing its own source lines. I look at how the reliability engineering works and why it matters more than the natural-language interface.

Nature published "Reimagining research papers as interactive and reliable AI agents" this week — Jiacheng Miao, Joe Davis, Yaohui Zhang, Jonathan Pritchard, and James Zou describing a system called Paper2Agent. The pitch is simple to state and harder to build well: take a paper plus its code repository and convert it into an AI agent that can run the paper's own methods on request, through natural language, instead of making the reader clone the repo and figure it out.
I want to walk through how it actually works, because the interesting part isn't the chat interface — it's the reliability engineering underneath it, which is the same problem I run into constantly building agentic systems for production: how do you let an LLM operate real tools without letting it also invent what those tools did.
The gap this closes
The paper opens with a problem every practicing scientist recognizes. A method paper is not usable on arrival. You have to find the repository, install the environment, work out the API surface, and map your data onto whatever object model the authors chose. The authors use AlphaGenome as their example: a capable genome-scale foundation model that nonetheless requires you to instantiate a client, hold an API key, construct variant objects correctly, and pick the right output modality before you get a single prediction. None of that is a criticism of AlphaGenome — it's just the standard tax that sits between a method existing and a method being used, and it falls hardest on people who are domain experts but not software engineers.
From paper to MCP server
Paper2Agent's answer is to represent the paper as a model context protocol (MCP) server — the same protocol that's become the standard way to expose tools and data sources to LLM agents, here pointed at a paper instead of a SaaS product. A pipeline of agents (Paper2MCP) reads the manuscript and codebase, sets up the execution environment, and wraps the paper's core analytical functions as MCP tools. An agent layer then wraps that server as a context provider, so any chat agent — the paper names Claude Code specifically — can connect to it and drive the paper's methods conversationally.

Each server has three parts, and the distinction between them is worth sitting with:
- Tools are the executable functions — the AlphaGenome tool that takes a variant and returns predicted effects on chromatin accessibility and expression, for instance. These ship with a pre-configured environment, so there's no dependency resolution left for the user.
- Resources are the static assets: manuscript text, code, datasets, figures, supplementary tables — stored in formats an agent can actually query rather than just read.
- Prompts encode the paper's multi-step workflows. A Scanpy prompt, for example, carries the correct sequence for preprocessing and clustering single-cell data, so the agent doesn't have to reconstruct analysis order from a methods section.
That three-way split is effectively the paper's own working memory, made machine-addressable: what it can do, what it knows, and how it proceeds.
Where the reliability actually comes from
The part I'd flag to anyone building agents against real tools: Paper2Agent does not let the LLM generate analysis code live for every query. Each tool is validated once against the reference codebase's reported results and figures, using the paper's own example data, and then locked. After that, invoking the tool runs the locked, tested code path — the model isn't regenerating the implementation each time it's called. The paper is explicit about why: this is meant to guard against "code hallucination," where LLM-generated code that looks plausible produces quietly wrong scientific output. Every tool also carries a reference back to the specific line in the original paper or codebase, so a result can be traced to its source rather than trusted on faith.
That's the actual innovation, in my read. Wrapping a paper's PDF text in retrieval-augmented generation is not new and doesn't get you reproducibility — it gets you a paper that answers questions about itself, not one that runs its own analysis correctly. Paper2Agent's claim to reliability rests on separating the one-time, verified act of building and testing a tool from the repeated, natural-language act of invoking it. The LLM's job at query time is routing and interpretation, not code generation.
What it looks like end to end
The authors validate the approach with three case studies. An AlphaGenome-based agent interprets genomic variants without the user touching API objects. Agents built on Scanpy and TISSUE (transcript imputation with spatial single-cell uncertainty estimation) run single-cell and spatial transcriptomics analyses and reproduce the original papers' results on the same example data. Most notably, they show multiple paper agents collaborating — composed together — to prioritize a causal gene for psoriasis, which is a preview of agents drawing on more than one paper's methods within a single analysis. Servers can be hosted remotely, on platforms such as Hugging Face Spaces, which also removes the local-environment problem entirely: nothing to install, because nothing runs on your machine.
My take
MCP already solved the integration problem for SaaS tools and internal APIs — a standard interface an agent can discover and call without bespoke glue code per service. Paper2Agent is that same move applied to the unit of scientific communication itself, and it's a sensible one, because a paper's method section is already, functionally, an API description written for humans. Turning it into one written for agents is a smaller leap than it first sounds.
What I'd watch is whether the validate-and-lock step scales to the messier majority of papers that don't ship clean, well-tested repositories the way AlphaGenome and Scanpy do — reproducibility here is bounded by how much of the original code was reproducible to begin with. But the design choice to separate verified execution from conversational interpretation is the right one, and it's the same lesson I keep relearning building agents in production: the value of an agent isn't that it can talk to your tools, it's that you can trust what it says the tools did.