OpenAI and V7: Giving AI Agents Institutional Memory with a Context Graph

·12 min read·Evergreen Tools Team

Today's models reason well, but they do not automatically understand the business context behind the task. Which fund report is current? How is the same entity named across three different systems? That context lives in documents, data rooms, spreadsheets, emails, and internal tools: scattered, unresolved, and invisible to agents. On September 21, 2026, OpenAI published a customer story about exactly this problem. V7's Go platform organizes that buried context into a Context Graph that agents can query and act on. The sentence worth remembering is not any single benchmark score but a line from the story: finance firms that get real value from AI will not be the ones with the most agents, but the ones with the best context.

Documents and data rooms

Context hides in documents, spreadsheets, and email, invisible to agents

1. The Real Cost: Rediscovering Context on Every Request

The default behavior in most agent architectures is stateless and re-queries everything. A user asks a question, and the agent searches SharePoint, then Drive, then email, then stuffs the hits into a prompt. In a demo this looks fine. In production it costs three things. First, every request repeats dozens of searches, so time and tokens scale linearly with request volume. Second, you only ever see text, never edges: the relationship between the fund in this CIM and the same fund in a custodian agreement from three years ago is never assembled. Third, and most insidious, the model relearns your business from scratch each time. Whatever understanding it built yesterday is gone today.

// The expensive default: every request rediscovers context.
async function naivelyAnswer(question) {
  const hits = [];
  hits.push(...await search("sharepoint", question));
  hits.push(...await search("drive", question));
  hits.push(...await search("email", question));   // dozens of searches...
  const prompt = hits.map((h) => h.text).join("\n");  // ...all stuffed into context
  return llm.ask(prompt, question);
}
// Cost grows with every request, and relationships between files are
// invisible because you only ever see text, never the edges.

2. The Context Graph: Making Relationships First-Class

V7 Go's answer is to structure the context first. When data arrives, it connects to repositories such as SharePoint and Google Drive, scans them for entities, relationships, facts, attributes, and metrics, and populates a graph. V7 says traversing that graph is an order of magnitude cheaper and faster than long-context approaches. Three design decisions stand out. Facts are bound to evidence, so every record preserves a citation back to the original file and every downstream step stays auditable. Graph queries coexist with a RAG fallback, so when the graph does not contain enough information the system still searches the underlying documents. And ingestion is incremental: when a new file lands, V7 identifies the companies, funds, or people in it and connects each fact to a new or existing record instead of rebuilding the graph. The deeper point is that a model can work with a firm's history without relearning it every time.

# Build a canonical key before you build a graph.
# "Acme Fund III LP", "ACME FUND III, L.P." and "Acme III" must be one node.
import re, unicodedata

SUFFIX = {"lp", "llc", "inc", "ltd", "fund", "iii", "ii", "i"}

def canonical(name: str) -> str:
    s = unicodedata.normalize("NFKD", name).lower()
    s = re.sub(r"[^a-z0-9 ]", " ", s)
    tokens = [t for t in s.split() if t and t not in SUFFIX]
    return " ".join(tokens)

print(canonical("Acme Fund III, L.P."))   # -> "acme"
print(canonical("ACME FUND III LP"))      # -> "acme"

3. The Numbers: From the HERB Benchmark to the Hardest Graph Queries

V7 has tested what structure is actually worth. On HERB, a benchmark for finding and connecting information spread across enterprise systems, V7's retrieval-only system outperformed the official baseline by 69% and reduced hallucinations on un-answerable queries by 38%. On the production path, it says agents complete 50 to 100 step workflows in minutes, reaching 99.9% accuracy while keeping an auditable trail of every decision. Model choice is tiered: high-volume structured extraction goes to GPT-5.6 Luna, reasoning and tool use go to GPT-5.6 Terra or Sol, and only the hardest graph queries go to GPT-6 Astra. On V7's own four-level graph-query set, GPT-5.6 Sol scored 78% at the very-hard level while GPT-6 Astra scored 89%; at the easy, medium, and hard levels both models were close to 100%. In other words, the frontier premium belongs on one narrow slice, not on every call. Moving document-heavy workloads from the Chat Completions API to the Responses API also reduced token use by roughly 5% on some PDF-heavy workflows and improved caching reliability.

// Populate the graph on arrival, and keep the evidence link.
async function ingest(file) {
  const extracted = await model.extract(file.text, ONTOLOGY);

  for (const fact of extracted.facts) {
    const node = await graph.upsertEntity({
      type: fact.entityType,
      key: canonical(fact.entityName),
      attributes: fact.attributes,
    });

    await graph.addFact(node.id, fact.predicate, fact.value, {
      sourceUrl: file.url,        // link every fact back to its origin
      page: fact.page,
      extractedAt: new Date().toISOString(),
      model: "gpt-5.6-luna",
    });
  }
}
// If the graph has no answer, fall back to RAG over the raw documents.

4. What the Results Look Like on the Business Side

These engineering decisions eventually turn into business numbers. Asset managers screen deals 21 times faster than before, collapsing a full-day process into fifteen minutes. One financial services team cut review time from more than 100 hours to under 10, saving $12,000 in expert costs per task. An insurance team reduced errors in claims processing by 13.5% against a manual baseline after giving its agents historical knowledge of previous claims and existing policies. The common thread is that none of these numbers came from swapping in a stronger model. They came from giving the model the right context.

# Route by difficulty instead of defaulting to the frontier model.
# V7's tiers: GPT-5.6 Luna for high-volume extraction, Terra/Sol for
# reasoning and tool use, GPT-6 Astra only for the hardest graph queries.
TIERS = [
    ("gpt-5.6-luna",  lambda q: q.kind == "extract"),
    ("gpt-5.6-terra", lambda q: q.kind == "reason"),
    ("gpt-6-astra",   lambda q: q.graphComplexity == "very_hard"),
]

def pick_model(query):
    for model, matches in TIERS:
        if matches(query):
            return model
    return "gpt-5.6-sol"

# On V7's hardest benchmark set, GPT-6 Astra scored 89% while the
# previous default scored 78% - a 11-point gain worth paying for
# on this narrow slice, and not on every call.

5. Building It Yourself: Five Reusable Steps

First, do entity resolution before you build a graph. Map Acme Fund III LP, ACME FUND III, L.P., and Acme III to the same key, or your graph fragments into disconnected points. Second, bind every fact to its evidence on write. Each fact must point back to the source file and page; in regulated industries, context without provenance is a liability. Third, ingest incrementally rather than rebuilding. Fourth, route by difficulty, reserving the frontier model for the narrowest hard slice. Fifth, manage a context budget: keep recent turns in active context and push older material into the graph for retrieval on demand. V7 also exposes Context Graph querying and ingestion through its MCP server, so customers can use it from ChatGPT and build workflows through MCP in Codex, which cut the time to create a medium-length workflow from about an hour to roughly twenty minutes.

# Context budget: keep recent turns live, push the archive to the graph.
MAX_LIVE_TURNS = 12
MAX_TOKENS = 60_000

async function assembleContext(session, question) {
  const recent = session.messages.slice(-MAX_LIVE_TURNS);
  const recalled = await graph.query(question, { limit: 40 }); // source-linked

  const ctx = [
    ...recalled.map((f) => `[${f.sourceUrl}] ${f.statement}`),
    ...recent.map((m) => `${m.role}: ${m.content}`),
  ];

  const tokens = countTokens(ctx);
  if (tokens > MAX_TOKENS) {
    // Drop the cheapest-utility evidence first, never the citation.
    return trimByUtility(ctx, MAX_TOKENS);
  }
  return ctx;
}

6. A Checklist and the Longer-Term Direction

Four checks. First, is your agent re-searching the same data on every request? If so, measure the wasted tokens before you do anything else. Second, does every fact carry a source link? Context without provenance is a liability in regulated work. Third, have you tiered model routing by difficulty, or does every request hit the most expensive model? Fourth, is your memory a queryable structure, or a pile of history nobody dares delete? V7 describes the longer-term goal as making shared memory proactive: when a fact in the graph changes, a workflow starts, flags inconsistencies, and shows people which analyses need another look. If a fund report is restated, the system surfaces the work still relying on the old figures. That is the difference between institutional memory and chat history. One tells you when something changed. The other just quietly goes stale.

Graph-shaped data relationships

Make relationships first-class and traversal gets an order of magnitude cheaper

Tiered model routing

Spend the frontier premium only on the narrowest hard slice

📌 Frequently Asked Questions

How is a Context Graph different from RAG?

RAG retrieves text one request at a time, while a Context Graph persists entities, relationships, and facts and keeps a citation to the original file for each fact. V7 reports traversing the graph is an order of magnitude cheaper and faster than long-context approaches, with RAG still available as a fallback.

Why is rediscovering context on every request so costly?

Because agents are stateless loops: each request repeats dozens of searches and re-stuffs hits into the prompt, so time and tokens scale with request volume, and the model relearns your business from scratch each time.

Where does the 89% figure come from?

On V7's own four-level graph-query set, GPT-5.6 Sol scored 78% at the very-hard level and GPT-6 Astra scored 89%; both were near 100% at the easy, medium, and hard levels, which shows the frontier gain is concentrated in the hardest slice.

Which model should handle which job?

Use GPT-5.6 Luna for high-volume structured extraction (78% lower cost per document than GPT-5.4 mini and 11.6 points higher accuracy), GPT-5.6 Terra or Sol for reasoning and tool use, and GPT-6 Astra only for the hardest graph queries.

What step do teams skip most often?

Entity resolution. Canonicalize the different names the same entity carries across systems before building the graph, or the graph shatters into disconnected points and cross-document relationships never come together.