Security2026-10-018 min read

OpenAI Lets You Run the Codex Harness on Your Own Machines. The Firewall Will Not Notice

The OpenAI Agents API, in public beta since September 2026, lets a managed Codex harness drive commands on a laptop, container or Lambda function you own through a codex exec-server that dials out over a WebSocket. That keeps code and data inside your network, and it also moves most of the security work onto you.

OpenAI opened its Agents API as a public beta in September 2026, and it went into DevDay on 29 September as the plumbing behind the Codex and Dots announcements. The API is a managed version of the Codex harness: OpenAI runs the loop that calls the model, compacts context, coordinates sub agents and keeps sessions alive for minutes or days, and you choose where the agent actually runs commands and touches files. The choices are an OpenAI hosted sandbox, one of the supported sandbox providers such as Cloudflare, E2B, Modal, Daytona or Vercel, or your own infrastructure. There is no separate Agents API fee during the beta, so you pay for model tokens and paid tools, plus whatever the environment costs you.

The self hosted option is the one that will appeal to regulated teams, and it is the one worth reading carefully. In that mode you run a small process, codex exec-server, on a machine you control and point it at a session. Developer write ups and the OpenAI documentation describe it as connecting outward to the harness over a WebSocket, after which the harness sends it commands to run and files to read or write. The examples OpenAI gives are a laptop, a Docker container and an AWS Lambda function. The selling point is obvious: your source code, build secrets and customer data never have to be uploaded to a cloud sandbox, because the work happens next to them.

Notice what the outbound connection means for your network controls. Most firewalls, security groups and change processes are built around inbound exposure. A process that dials out to an OpenAI endpoint on port 443 looks like any other HTTPS client, needs no inbound rule, no load balancer and no ticket, and once connected it is a channel through which a remote service decides what runs on that host. That is a legitimate architecture, and it is the same shape as remote management agents and CI runners. It is also exactly the shape your detection rules are least likely to flag, which is why we argued in our ClosedQuorum piece that LLM API traffic has become an egress evidence problem rather than a curiosity.

The responsibility split also changes. In the hosted sandbox, OpenAI or the sandbox provider owns isolation, network policy and teardown. In self hosted mode, OpenAI describes you as owning provisioning, reconnection and shutdown, and community projects evaluating the exec-server have listed the rest of the job: choosing the working directory, isolating each task in its own Git worktree, denying access to credential files, keeping outbound network access off unless a task needs it, gating risky actions with host side approvals, and confirming that child processes are cleaned up when a session ends. None of that is enforced by the API on your behalf. If an engineer runs the exec-server from their normal shell on their normal laptop, the agent inherits everything that shell can reach.

The API key is the second boundary. OpenAI guidance is to keep the key outside the agent environment and to create it with narrow scopes, naming api.agents.read, api.agents.write and api.responses.write. Follow that literally. A key with full project scope sitting in an environment variable inside the same container the agent controls is a key the agent, or anything that compromises the agent through a poisoned repository, can read and reuse. Our GitSpawn article showed how little it takes for a repository to become the payload, and a self hosted executor is the place where that payload runs with your network position rather than a disposable sandbox.

Map it to the frameworks you already report against. Under ISO 27001, Annex A 8.20 to 8.22 on network security and segregation, 8.9 on configuration management, 5.23 on cloud services and 8.15 and 8.16 on logging and monitoring all apply to a host that takes instructions from a third party service. For SOC 2, CC6.6 covers boundary protection against threats from outside the system, CC6.8 covers preventing unauthorised software, and CC7.2 expects you to detect anomalous activity, which you can only evidence if the commands an agent ran are logged on your side, not only in an OpenAI dashboard. ISO 42001 adds the AI register and human oversight expectations: each self hosted executor is an AI system deployment with an owner, a purpose and a defined approval rule. If you run Vanta, Drata or Secureframe, add executor hosts as assets with their own control owners rather than folding them into the OpenAI vendor record.

The practical controls are not exotic. Run executors only in dedicated, short lived containers or virtual machines with a read only base image, a scratch workspace and no mounted home directory. Give them an egress allowlist that permits the OpenAI endpoint and the package mirrors a task needs, and nothing else, so a compromised session cannot reach your internal network or exfiltrate to arbitrary hosts. Inject secrets per task through your secrets manager with short expiry, and never let the agent see the API key that authorises the session. Log every command and file write locally to your SIEM. Then look for executors you did not deploy: search endpoint telemetry for codex exec-server processes and long lived WebSocket connections to OpenAI from developer machines and build agents.

A short list for this week. Decide whether self hosted Codex execution is allowed at all, and if so, write down where, as a line in your AI acceptable use policy. Publish an approved executor image so engineers have a safe default rather than their laptop. Rotate any OpenAI keys that were created with broad scope during early testing and reissue them with the three agent scopes only. Add the executor pattern to your threat model and your secure development policy, and record each deployment in your AI register. Our policy templates include an AI acceptable use policy and a secure development policy you can extend with these clauses, and the compliance readiness checklist will show where agent logging evidence is missing.

Our view is that self hosting the executor is the right choice for most companies with sensitive code, and it is better than shipping a monorepo and production secrets into a sandbox someone else runs. But keeping data inside the network is not the same as keeping control inside it. The model and the harness still decide what happens next. Self hosted mode is only safer than the hosted sandbox if your own isolation is at least as good as theirs, and on day one, for most teams, it will not be.

OpenAICodexAgents APIexec-serverself-hosted agentscoding agentsegress controlshared responsibilityISO 27001ISO 42001SOC 2VantaDrata

Editorial note: AES Tech reviews are independent. Some outbound links are affiliate links and are marked sponsored; they never change our rankings. See our disclosure.

// Signal, not noise

Get the next post by email

One short email when something worth knowing ships. No spam, unsubscribe anytime.

Loading comments...

Add a comment

Corrections and first-hand experience are the most useful things you can leave. Comments are screened automatically and reviewed by a human; see the moderation policy.

0/4000 · plain text · links are held for review

More from the blog