The Codex Harness as a Service: Inside OpenAI's New Agents API
💡 Tool Tip:JSON Formatter, Env File Generator, API Response Time Calculator
On September 10, 2026, OpenAI put the harness behind Codex behind an API. The Agents API is in public beta, and the pitch is narrow and concrete: describe a task, a model, a set of tools, and an environment, and OpenAI runs the agent loop for you. That loop - context compaction, tool selection, subagent coordination, sessions that stay alive for hours - is exactly the part most teams have spent two years rebuilding badly. So it is worth being precise about what you are buying, and about the responsibilities that do not transfer with it.
An agent loop you no longer have to write yourself
1. A Session, Three Kinds of Compute
The unit of work is a session. You create one with the task text, the model, the tool list, and the environment, and you receive events and output back. OpenAI hosts and maintains the harness; the compute decision stays yours. There are three options: an OpenAI-hosted sandbox that runs on the same infrastructure that powers Codex and ChatGPT, your own infrastructure, or a partner sandbox - Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel are the named integrations. Hosted sandboxes are provisioned and managed by OpenAI and can be configured with your files, packages, skills, and plugins. Secrets are referenced through vault identifiers rather than inlined. OpenAI states there are no additional fees for using the API itself: you pay for the tokens and tools your agents consume.
// 1. A production-shaped agent in one call (Agents API, public beta, Sep 10 2026)
import OpenAI from "openai";
const client = new OpenAI();
const session = await client.beta.agents.sessions.create({
agent: {
model: "gpt-6-astra",
tools: [
{
type: "mcp",
server_label: "observability",
transport: { type: "http", server_url: "https://observability.example.com/mcp" },
},
],
multi_agent: { enabled: true, max_concurrent_subagents: 3 },
},
vault_ids: ["vault_YOUR_VAULT_ID"],
environment: {
type: "openai_hosted",
capability_directories: ["/workspace/capabilities/skills"],
},
input:
"Investigate service-api elevated 5xx rate over the last 30 minutes. " +
"Delegate deployment, error, and dependency analysis to subagents. " +
"Save findings, evidence, and recommended mitigation in /workspace/outputs.",
});2. What the Harness Takes Over
Four harness behaviours are worth the migration on their own. First, context management: as a session approaches its context limit, the API automatically compacts earlier context, so multi-window workflows no longer need your own summarisation logic. Second, tool search: tool definitions are loaded on demand instead of pasted into every prompt, which cuts token usage and, importantly, preserves the upstream prompt cache. Third, programmatic tool calling: agents can run calls in parallel, chain related operations, and filter or join results in code, so only the reduced result returns to the context window. Fourth, a single tool surface: MCP servers, custom functions, and built-in tools such as web search are declared in the same tools array. OpenAI also describes the harness as versioned - it improves alongside model launches - so you inherit those gains without rewriting the loop yourself.
# 2. Tool search loads definitions on demand, so the prompt cache stays warm
agent_tools = [
{"type": "tool_search", "index": "internal-tools"},
{"type": "programmatic_tool_calling"}, # loop, join and filter in code
{
"type": "mcp",
"server_label": "openai_docs",
"transport": {"type": "http", "server_url": "https://developers.openai.com/mcp"},
},
]
run = client.beta.agents.sessions.create(
agent={"model": "gpt-6-astra", "tools": agent_tools},
input="Summarise every failed deployment in the last 24 hours, grouped by service.",
)
# Only the reduced result returns to the context window, not 400 raw tool rows.Sandboxes and egress rules are infrastructure decisions, not prompts
3. Subagents Without an Orchestrator of Your Own
Multi-agent support is a flag, not a framework. Set multi_agent.enabled and a concurrency cap, and the API splits a task into independent pieces, gives each subagent its own context, and has the coordinating agent assemble the results. Each subagent holding separate context is the point: one focused context window drifts less than a single window juggling three investigations. The operational caveat is arithmetic. The published example caps concurrency at three subagents, which is the fan-out; the cost is three concurrent streams of spend against the same budget. Delegate when the work is genuinely separable, and measure the token difference before you leave the flag on by default.
// 3. Subagents keep their own context; the coordinator keeps the plan
const session = await client.beta.agents.sessions.create({
agent: {
model: "gpt-6-astra",
multi_agent: { enabled: true, max_concurrent_subagents: 3 },
},
input: "Research three vendors, one subagent each, then write a comparison table.",
});
// Fan out only for work that genuinely needs separate context:
// three concurrent subagents are three concurrent streams of token spend.4. An Open-Source Substrate, Operated For You
The harness is the open-source Codex harness. The codebase is public, so the coordination logic between model calls, tools, and context is inspectable - useful when you have to explain an agent's behaviour to a reviewer, a customer, or an auditor. The more accurate reading of the announcement is a split: OpenAI operates and maintains the loop, and you keep the domain parts - tools, knowledge, and workflows. If your differentiation was an orchestration loop you wrote yourself, that layer is being commoditised. If your differentiation is the tool surface and the data behind it, nothing in this announcement threatens it.
# 4. Bring your own compute: same harness, your network, your compliance rules
environment:
type: self_hosted
runner: fixed # or on_demand
provider: modal # Blaxel, Cloudflare, Daytona, DigitalOcean,
# E2B, Modal, Oracle, Runloop, Vercel
egress:
default: deny
allow:
- observability.internal:443
- github.com:443
secrets: vault # never bake keys into the imageParallel subagents multiply cost as reliably as they multiply speed
5. What Does Not Transfer With the API
Five responsibilities stay with you. Cost: you pay per token and per tool call, so per-session budgets and compaction policy are your controls. Egress: a sandbox that can reach the internet is an exfiltration path, so default-deny egress and allow lists belong in the environment definition. Secrets: vault references, rotated, never baked into images or prompts. Idempotency: long-running sessions retry, so tool calls need idempotency keys. Observability: the event stream is your audit log, which means you choose the sink and the retention period. Data residency is the sixth: a hosted sandbox is convenient, and it is also someone else's region. The API is in public beta and OpenAI says it will iterate quickly, so pin the harness version and treat upgrades like any other dependency change.
{
"agent_guardrails": {
"harness_version": "pin-explicitly",
"session_budget_usd": 12.0,
"compaction": { "mode": "automatic", "keep_last_turns": 8 },
"tool_calls": { "idempotency_key": "session_id + tool_name + args_hash" },
"events": { "sink": "otel", "retain_days": 30 },
"environments": { "hosted": "no_pii", "self_hosted": "pii_ok" }
}
}📌 Frequently Asked Questions
What is the OpenAI Agents API?
A public beta announced on September 10, 2026 (OpenAI's announcement). It exposes the harness and infrastructure behind Codex through a single API call: specify the task, model, tools, and environment to create a session. Compute can be an OpenAI-hosted sandbox, your own infrastructure, or a partner sandbox.
Does the Agents API cost extra?
OpenAI states there are no additional fees for using the Agents API; you pay for the tokens and tools your agents use, as outlined on its pricing page.
Which sandbox partners are supported?
The named integrations are Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel, alongside the OpenAI hosted sandbox that shares infrastructure with Codex and ChatGPT.
How do subagents work?
Set multi_agent.enabled and max_concurrent_subagents in the agent configuration (the published example uses 3). The API breaks the task into independent pieces, each subagent keeps its own context, and the coordinating agent assembles the results. Cost scales with concurrency.
Is the harness open source?
Yes. OpenAI says the Agents API is powered by the open-source Codex harness, whose codebase is public at github.com/openai/codex. OpenAI operates and maintains it while developers can inspect the coordination logic.
🔧 Recommended Tools
JSON Formatter
Read the session and event payloads without squinting
AI Token Counter
Estimate the context you are actually paying for
Env File Generator
Keep sandbox and runner configs reproducible
API Response Time Calculator
Budget latency per session, not per call
Cron Expression Generator
Schedule long-running agent jobs predictably
📚 Sources
- OpenAI - Introducing the Agents API (September 10, 2026): managed Codex harness, hosted/self-hosted/partner sandboxes, no additional fees
- OpenAI - Agents API overview (developer documentation)
- OpenAI - Context compaction guide (automatic context management for long sessions)
- GitHub - openai/codex: the open-source harness that powers the Agents API