Stanford's Paper2Agent in Nature: Turn a Paper and Its Codebase Into an MCP Server

·11 min read·Evergreen Tools Team

Reproducing a computational paper is an expensive ritual: clone the repository, install dependencies, debug version conflicts, read half a README, and discover the authors' environment no longer runs. That cost keeps useful methods locked inside PDFs forever. A Stanford team (Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard, and James Zou) published Paper2Agent in Nature on September 16, 2026, with a different answer: do not change the paper, change the interface. Package the paper and its codebase as a Model Context Protocol server, so any agent can call the paper's methods through natural language. The authors describe the result as a virtual corresponding author. The framing matters more than the mechanics. For essentially all of human history, knowledge has been stored as a passive artifact: carved into stone, then typed into pages. Turning a paper into a server is an attempt to make the artifact answer back.

"Research papers and data"

"Turning a paper into callable tools"

1. What Paper2Agent Actually Does

According to Nature's reporting, Paper2Agent starts by accessing a paper's main text, code, datasets, and other elements, and deposits that material on an MCP server. A team of AI agents then autonomously writes tools that apply the paper's methods to fresh data, and any MCP-compatible client, Claude Code included, can call them in natural language. The code is MIT-licensed and installs as a skill for Claude Code or Codex. Prebuilt AlphaGenome, Scanpy, and TISSUE servers run on Hugging Face Spaces, and there is a hosted version at paper2agent.ai.

# 1) Pin the environment first. Parsing is worthless if it will not run.
FROM python:3.12-slim
RUN pip install --no-cache-dir uv
WORKDIR /paper
COPY pyproject.toml uv.lock ./
RUN uv sync --frozen --no-dev          # lockfile, reproducible
COPY . .
RUN python -m pytest tests/test_repro.py -q   # must pass before tooling

2. The Pipeline: From PDF to Callable Tools

The pipeline splits into four steps. First, parsing: extract the narrative, the intent of the equations, and the runtime requirements. Second, assembly: place the companion code into a containerized environment so it genuinely runs with its dependencies. Third, tooling: agents break the methods into schema-carrying tools exposed through the MCP tool interface, complete with argument validation. Fourth, verification: regress against the paper's original results to confirm behavior did not drift during toolification. Step three deserves attention, because the tools are generated by agents reading code rather than written by hand. That is the source of the speed and the main reliability risk.

# 2) Expose a paper method as an MCP tool with a validated schema.
from mcp.server.fastmcp import FastMCP
from pydantic import BaseModel, Field

mcp = FastMCP("paper2agent-alphagenome")

class VariantQuery(BaseModel):
    sequence: str = Field(min_length=1, max_length=4096)
    assay: str = Field(default="ATAC-seq")

@mcp.tool()
def predict_variant_effect(q: VariantQuery) -> dict:
    """Apply the paper's method to new sequence input."""
    out = run_paper_method(q.sequence, q.assay)   # the original code path
    return {"score": out.score, "assay": q.assay, "tool_version": TOOL_VERSION}

if __name__ == "__main__":
    mcp.run()

3. The AlphaGenome Case: 45 Minutes, $14, 91.2%

The team validated the framework end to end on AlphaGenome, a model that predicts properties of DNA sequences. Building a working agent took about 45 minutes and roughly $14 of compute. It answered genetics questions with near-perfect accuracy and outperformed Biomni, a comparison tool. The overall prototype averaged 91.2% ± 1.6% accuracy on benchmark questions drawn from 100 computational biology papers. The interesting part is not that AI performed well. It is that reproduction cost dropped to the double-digit-dollar range, which changes not the ceiling of what a researcher can do but the size of the candidate set that is worth reproducing. The cost figure deserves a second look, because it reframes the incentive. At roughly $14 per reproduction, trying a method you are merely curious about becomes rational, and the bottleneck moves from compute to deciding which papers are worth asking about.

# 3) One regression case per tool, expected value from the paper itself.
REGRESSION = [
    # (input fixture, paper-reported expectation, tolerance)
    ("fixtures/alphagenome_case_01.json", {"score": 0.812}, 1e-3),
    ("fixtures/alphagenome_case_02.json", {"score": 0.447}, 1e-3),
]

def verify_tools(call_tool):
    failures = []
    for fixture, expected, tol in REGRESSION:
        got = call_tool(load(fixture))
        for k, v in expected.items():
            if abs(got[k] - v) > tol:
                failures.append((fixture, k, v, got[k]))
    return failures

print(verify_tools(call))   # empty list means no drift

4. The Pattern Worth Stealing, Not Just the Tool

Swap the paper for something else and the pattern holds. The legacy service nobody on your team dares to touch is a paper without a body: the code exists, the data exists, but the behavior logic lives only in a few veterans' heads. Apply the same approach, parse, containerize, toolify, verify, and you get a service you can ask questions of. New agents can query it and call it in natural language instead of reading three years of code first. Paper2Agent's real contribution is a domain-agnostic recipe: parse, assemble, toolify, verify.

// 4) Bind generated tool versions to a source commit. Rebuild on change.
const MANIFEST = {
  paper: "alphagenome-2026",
  sourceRepo: "https://github.com/example/alphagenome",
  sourceCommit: "9f3c1ab",       // the exact revision the tools were built from
  toolVersion: "1.0.0",
  builtAt: "2026-09-22T08:00:00Z",
};

async function needsRebuild(manifest) {
  const head = await getHeadCommit(manifest.sourceRepo);
  if (head !== manifest.sourceCommit) {
    console.warn("source moved from", manifest.sourceCommit, "to", head);
    return true;                  // regenerate tools, then re-run regression
  }
  return false;
}

5. Where the Risk Lives: Reliability and Provenance

Auto-generated tools bring three classes of problem. The first is silent drift: the wrapper the agent wrote looks correct but diverges from the original implementation on edge cases the benchmark does not cover. The second is version drift: when the paper's code updates, the generated tools are not rebuilt in lockstep. The third is provenance: when a conclusion comes from an agent calling a tool another agent generated, the audit chain is much longer than with a plain script, and failures are harder to localize. The practical answer is to treat regression tests, version pinning, and call tracing as required components of the pipeline rather than options. None of these risks is exotic, and all three are manageable with ordinary software discipline. What makes them dangerous is the speed of generation: you can build thirty tools in an afternoon and have no tests for any of them.

// 5) Trace every agent-to-tool call. Long audit chains are hard to debug.
function tracedToolCall(agentId, toolName, args, fn) {
  const traceId = crypto.randomUUID();
  const started = Date.now();
  try {
    const result = fn(args);
    log({ traceId, agentId, toolName, args, ok: true, ms: Date.now() - started });
    return result;
  } catch (err) {
    log({ traceId, agentId, toolName, args, ok: false, error: String(err) });
    throw err;
  }
}

export const predict = (a) =>
  tracedToolCall("research-agent", "predict_variant_effect", a, callTool);

6. A Getting-Started Checklist

Five rules. First, pin the environment with an image and a lockfile, or nothing runs no matter how well you parsed. Second, write one minimal regression case per generated tool, using the paper's original results as the expected value. Third, bind tool versions to the source repository commit hash and rebuild when the source moves. Fourth, log the agent call chain so a failure can be traced to a specific tool and argument set. Fifth, try it on an internal repository before production. This approach turns reproduction from a one-off human effort into a repeatable pipeline, and a pipeline is valuable precisely because it gets run many times.

"Code and lab environment"

"Reproduction cost down to double-digit dollars"

"Compute and automation"

"Regression tests, version pinning, call tracing"

📌 Frequently Asked Questions

Who built Paper2Agent and where was it published?

It comes from a Stanford team: Jiacheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard, and James Zou. It was published in Nature on September 16, 2026.

How is it different from ordinary paper companion code?

Companion code expects a human to read it, install it, and debug it. Paper2Agent wraps the paper and its codebase as an MCP server, so any MCP-compatible agent can call the methods directly in natural language.

What did the evaluation show?

On AlphaGenome, a working agent was built in about 45 minutes for roughly $14 of compute, answering genetics questions with near-perfect accuracy and beating the comparison tool Biomni. The prototype averaged 91.2% ± 1.6% on benchmark questions from 100 computational biology papers.

Can I use it today?

Yes. The code is MIT-licensed and installs as a skill for Claude Code or Codex. Prebuilt AlphaGenome, Scanpy, and TISSUE servers run on Hugging Face Spaces, with a hosted version at paper2agent.ai.

What is the biggest risk?

Auto-generated tools can drift silently on edge cases, go stale when the source repository updates, and lengthen the audit chain. The countermeasures are regression tests, version pinning, and call tracing.