Back to all articlesSystem Integration

The Agent Sandbox Is a Security Boundary: Files, Egress, Credentials, and Escape Tests

Containers alone are not an agent security model. Build layered isolation with default-deny policy, brokered credentials, controlled egress, and escape tests.

Mapki

Mapki

Oct 9, 2026•10 min read
Share:
The Agent Sandbox Is a Security Boundary: Files, Egress, Credentials, and Escape Tests

The Agent Sandbox Is a Security Boundary: Files, Egress, Credentials, and Escape Tests

An agent that can execute code, open files, call tools, or browse the network is not just generating text. It is operating a runtime.

That runtime needs a security boundary independent of the model’s intentions. Prompt instructions can be forgotten, poisoned, or misinterpreted. A sandbox turns the most important limits into enforceable controls around the process.

The common mistake is to equate containerized with sandboxed. A container is one isolation mechanism. It is not automatically a complete policy, identity, network, filesystem, or escape-resistance model.

Start with the actual trust boundary

Write down what the agent may execute and what it must never reach. A useful threat model includes five surfaces:

  • Code execution: Generated scripts, package installation, shell commands, interpreters, and local MCP servers.
  • Filesystem: Workspace files, mounted host paths, credentials, temporary files, and persistent project state.
  • Network: Package registries, APIs, metadata endpoints, internal services, arbitrary internet destinations, and DNS.
  • Authority: API keys, cloud roles, service accounts, browser sessions, and tool permissions.
  • Resources: CPU, memory, disk, processes, file descriptors, bandwidth, and execution time.

The model may be unreliable on any of these surfaces. Treat generated code, tool output, repositories, documents, and MCP servers as untrusted inputs.

SandboxEscapeBench frames the boundary precisely: its evaluated trust boundary is container-to-host isolation, and successful escape means reading a host file outside the container namespaces. Its 18 tasks span orchestration, runtime, and host/kernel layers. The benchmark excludes hypervisor escapes and hardware side channels, so passing a container benchmark does not prove that a production environment is safe. It proves something narrower: that the tested controls resisted the tested classes of attack.

A layered runtime architecture

Use multiple controls because each layer fails differently.

Agent process
|
| untrusted shell / Python / local MCP request
v
Policy broker outside the agent process
|  identity + tool args + event history + audit
v
Egress gateway ---- short-lived credential injection
|  default deny, DNS rules, host/path allowlists
v
Sandbox boundary
|  microVM, gVisor, or hardened container
v
Ephemeral workspace + bounded persistent volume
|
v
Approved external systems, never the host network by default
  • Process boundary: Prevent direct agent access to the policy engine and host control plane.
  • Filesystem boundary: Expose only a dedicated workspace. Avoid host mounts and Docker socket access.
  • Network boundary: Route all outbound traffic through a gateway that can inspect destination, method, path, identity, and policy state.
  • Credential boundary: Keep secrets outside the sandbox. Inject them only into approved requests.
  • Runtime boundary: Choose containers, gVisor, or microVMs based on the threat model and tenancy.

Strands Box is a useful open-source reference for this shape. Its OS restrictions live in box.toml, while semantic policies are written in Dogwood. Shell, Python, the egress gateway, and the MCP broker consult an engine outside the agent process. Its default behavior is deny unless a policy explicitly permits an operation, and it records policy decisions as OTLP JSON.

Choose the isolation tier deliberately

There is no universal “best sandbox.” The right choice depends on whether the code is trusted, whether tenants share infrastructure, and whether kernel escape is in scope.

  • Hardened container: Fast and dense. Use only for reviewed code or low-risk, single-tenant workloads. Drop capabilities, apply seccomp and AppArmor, make the root filesystem read-only, and never expose a host socket.
  • gVisor: Adds a user-space kernel and intercepts syscalls before they reach the host kernel. It generally provides stronger isolation than a standard container with lower startup cost than a VM, with some I/O overhead.
  • Firecracker microVM: Gives each workload a separate guest kernel and a hardware-assisted boundary. It is appropriate for untrusted code, multi-tenant execution, and higher-consequence workloads.
  • Kata Containers: Provides VM-backed isolation behind container-oriented orchestration. It is useful when Kubernetes workflows are required but the shared-kernel boundary is not sufficient.

The choice is not just a performance comparison. A fast container that exposes /var/run/docker.sock is less safe than a slower microVM with controlled egress and no host mounts.

Default-deny policy should cover behavior, not only paths

A path allowlist is necessary but incomplete. An agent can read an allowed file, infer a secret, and send it through an allowed-looking network request. Policies should be able to consider the operation, arguments, identity, earlier actions, and elapsed time.

sandbox:
name: invoice-analysis
root_fs: read_only
workspace:
path: /workspace
persistent: true
max_bytes: 536870912
resources:
cpu_millis: 1500
memory_mb: 2048
pids_max: 128
wall_time_seconds: 300
outbound_bytes_max: 10485760

filesystem:
allow:
- path: /workspace
read: true
write: true
- path: /tmp
read: true
write: true
deny:
- /etc/shadow
- /var/run/docker.sock
- /proc/kcore
- /sys

network:
default: deny
dns:
allow: [api.openai.com, pypi.org, files.pythonhosted.org]
https:
allow:
- host: api.openai.com
methods: [POST]
paths: [/v1/responses]
- host: files.pythonhosted.org
methods: [GET]
metadata_endpoints: deny

credentials:
inject_via: egress_gateway
agent_visibility: none
max_ttl_seconds: 300

execution:
commands: [python, node, git]
package_install: deny
local_mcp_servers: deny
require_approval_for: [network_expansion, persistent_write, publish]

The policy is a starting point, not proof of safety. Test the enforcement path. If a rule is implemented only in the agent’s prompt, it is guidance. If it is enforced by the broker, gateway, kernel, or hypervisor, it is a control.

Keep credentials out of the workspace

Never place a cloud key, database password, or browser cookie in an environment variable that arbitrary generated code can read. The agent should request a capability, not receive the underlying secret.

type EgressRequest = {
sandboxId: string;
host: string;
method: string;
path: string;
capability: string;
};

async function forward(req: EgressRequest, body: Uint8Array) {
const decision = await policyEngine.check({
subject: req.sandboxId,
operation: "http_request",
resource: `${req.method} https://${req.host}${req.path}`,
capability: req.capability,
});

if (decision.effect !== "allow") {
throw new Error("egress denied");
}

const token = await credentialBroker.issue({
capability: req.capability,
audience: req.host,
ttlSeconds: 300,
});

return httpClient.request({
host: req.host,
method: req.method,
path: req.path,
headers: { authorization: `Bearer ${token}` },
body,
});
}

The response should be filtered before returning to the agent. Remove secrets, unrelated tenant data, and unnecessary headers. Log the decision and a safe request hash, not the credential or raw sensitive payload.

Treat MCP servers and tool calls as code execution

A local MCP server is a program. It can read files, invoke a shell, load packages, or make network requests. Do not grant a server broader access than the agent itself.

For every tool call, enforce:

  1. Caller identity: Which agent, tenant, and workflow requested it?
  2. Tool version: Which server build and schema are active?
  3. Argument policy: Are paths, identifiers, amounts, and destinations allowed?
  4. Budget: Does the call fit the run’s resource and spend envelope?
  5. Approval: Is a human or second control required?
  6. Audit: What was proposed, what was allowed, and what actually happened?

This complements an MCP gateway. The gateway controls the tool boundary; the sandbox controls the runtime that hosts code and local servers. They are different layers.

Preserve state without preserving authority

Long-running agents need persistence. A developer expects a project workspace to survive a pause, restart, or handoff. Persisting the entire machine, however, can preserve stale credentials, poisoned packages, or unauthorized process state.

Persist only an explicit workspace volume. Recreate the runtime and credentials for each session.

persistence:
save:
- /workspace/src
- /workspace/tests
- /workspace/results
- /workspace/manifest.json
discard:
- /tmp
- /run
- /workspace/.credentials
- process_state
- shell_history
restore_checks:
- verify_manifest_hash
- rescan_dependencies
- reapply_current_policy
- reissue_short_lived_credentials

A restore operation should be treated as a new admission decision. Verify the workspace, reapply the current policy, and issue new credentials. Do not resume a paused sandbox with yesterday’s authority.

Test the boundary, not just the happy path

Sandbox testing should include both ordinary policy tests and adversarial escape tests.

  • Filesystem tests: Attempt host path traversal, symlink escapes, device access, /proc inspection, and socket access.
  • Network tests: Attempt direct IP access, alternate DNS, redirects, IPv6, loopback, metadata endpoints, and exfiltration through permitted hosts.
  • Credential tests: Search environment variables, process arguments, mounted files, logs, core dumps, and error messages.
  • Tool tests: Try unapproved tools, schema drift, argument smuggling, local MCP startup, and cross-tenant identifiers.
  • Resource tests: Trigger fork bombs, memory growth, disk exhaustion, subprocess storms, and long sleeps.
  • Escape tests: Run a controlled subset of known container and runtime escape classes in a disposable environment. Never test an escape against production.
# Safe admission checks for a disposable sandbox
sandboxctl run --policy policy.yaml -- python - <<'PY'
import os
from pathlib import Path

assert not Path('/var/run/docker.sock').exists()
assert 'AWS_SECRET_ACCESS_KEY' not in os.environ
assert not Path('/etc/shadow').readable()
print('baseline checks passed')
PY

# Verify the egress gateway, not the agent, owns the credential
sandboxctl run --policy policy.yaml -- curl -fsS https://api.openai.com/v1/models
sandboxctl run --policy policy.yaml -- curl -fsS http://169.254.169.254/latest/meta-data/ || true

SandboxEscapeBench is useful as a model for the test philosophy: define the boundary, select reproducible vulnerability classes, report what is in and out of scope, and keep the evaluation environment isolated. Do not turn a benchmark into an exploit cookbook for a live system.

What practitioners are learning

A LocalLLaMA discussion frames the practical tradeoff clearly: Docker is familiar, microVMs provide stronger isolation, WASM is lightweight but capability-limited, and teams still need to decide how to support network access and persistent files. Those are product constraints, not reasons to run generated code on the host.

The LangSmith Sandboxes walkthrough describes the operational side of the problem: user-facing agents need fast startup and burst scaling, but also need to survive prompt injection, malicious MCP servers, container escapes, and pause/resume workflows. Its emphasis on an authentication proxy and network allowlists is a useful reminder that runtime isolation alone does not control what a permitted process can exfiltrate.

A Hacker News production discussion makes a related architectural point: observability and governance should sit in an independent execution layer between the agent and business systems. The agent proposes an intent; the execution layer verifies authority, budget, and policy before an external side effect.

References & Community Insights

Final checklist

Before allowing an agent to execute code or call external tools, verify:

  1. Is the agent runtime isolated from the host kernel, host filesystem, and host control sockets?
  2. Are filesystem and network permissions default-deny and scoped to the workflow?
  3. Are credentials brokered, short-lived, audience-bound, and invisible to generated code?
  4. Are local MCP servers and tool calls subject to the same policy as shell and Python?

5. Can persistent workspaces survive without preserving processes or authority?

  1. Are resource exhaustion, exfiltration, and known escape classes tested in disposable environments?
  2. Does an independent execution layer record intent, policy decision, and effect?

The safest agent is not the one with the most confident prompt. It is the one whose runtime makes unsafe behavior difficult, observable, bounded, and recoverable even when the model, tool, document, or dependency is compromised.

Want to implement this in your business?

Mapki designs bespoke AI agents, custom workflow automations, and tool-agnostic integrations tailored specifically to your existing ERP, CRM, and databases.

Tags:AI Agent SecuritySandboxingMicroVMZero TrustMCP SecurityAgent Runtime