mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

Security

Spens: Observable Sandboxes for Coding Agents

Docker, nono, and mitmproxy wired together to capture LLM calls, tool invocations, and network traffic from autonomous coding agents.

Source: spens.refwd.ai
Spens: Observable Sandboxes for Coding Agents

Coding agents need two things: isolation strong enough to contain mistakes and observability deep enough to debug them. Spens combines Docker containers, the nono sandboxing layer, and mitmproxy to create a reproducible execution environment that logs every LLM API call, tool invocation, HTTP request, and file change.

This is not a security product. It is a dev-tools stack that makes agent behavior visible and auditable. You run an agent in a container, intercept its network traffic, and review the session log afterward. If you need to understand what your agent did or replay a session with different parameters, Spens gives you the plumbing.

Architecture

Spens wraps four components into a single execution flow:

  • Docker container: Provides process and filesystem isolation. The agent runs as an unprivileged user inside a container that cannot access your home directory or system files.
  • nono: Adds a second sandboxing layer inside the container. Blocks syscalls and restricts capabilities beyond what Docker provides.
  • mitmproxy: Intercepts all HTTP and HTTPS traffic. Captures request and response bodies, headers, and timing. Swaps placeholder API keys for real credentials only when the request targets an approved domain.
  • Session logger: Records LLM API calls, tool invocations, file changes, and network activity in a structured log. After the session ends, a local web viewer lets you step through the timeline.

The flow looks like this:

  1. You run spens node-24 pi . to start a session with the Pi agent in a Node.js 24 runtime.
  2. Spens launches a Docker container with nono configured, mounts your workspace read-only, and starts mitmproxy.
  3. The agent makes an LLM API call. mitmproxy intercepts it, swaps the placeholder key for your real key, and logs the request and response.
  4. The agent invokes a tool (file write, shell command). Spens logs the invocation and result.
  5. The agent makes an HTTP request to an external service. mitmproxy checks the domain allowlist, blocks or allows the request, and logs it.
  6. The session ends. Spens writes the log to disk and opens the web viewer.

Network Interception

mitmproxy sits between the agent and the network. It terminates TLS connections, inspects traffic, and re-encrypts outbound requests. This breaks certificate validation unless you install mitmproxy’s CA certificate in the container, which Spens does during container setup.

The agent sees mitmproxy as a transparent proxy. It does not need to configure proxy settings or trust a custom CA. The container’s environment variables point all HTTP and HTTPS traffic to mitmproxy’s listening port.

Domain Allowlisting

You configure domain rules in a spens.yaml file:

domains:
  - host: api.anthropic.com
    methods: [POST]
  - host: api.openai.com
    methods: [POST]
  - host: pypi.org
    methods: [GET]

mitmproxy checks every request against this list. If the host is not allowed, the request fails with a 403. If the host is allowed but the method is not, the request fails. Non-HTTP protocols (raw TCP, UDP) are always blocked.

This is not a firewall. It is a policy layer that makes agent network behavior explicit. You decide which APIs the agent may call and which HTTP methods it may use.

Secret Injection

API keys never enter the container. You store them in your local environment or a secrets manager. Spens replaces them with placeholders in the agent’s environment variables:

ANTHROPIC_API_KEY=spens_placeholder_anthropic

When mitmproxy sees a request to api.anthropic.com, it swaps the placeholder for your real key in the Authorization header. The agent never sees the real key. If the agent leaks its environment variables, it leaks placeholders.

This works because mitmproxy inspects every request before forwarding it. The swap happens in memory, not on disk.

Sandboxing Trade-offs

Docker provides process isolation, filesystem isolation, and network isolation. nono adds syscall filtering and capability restrictions. Together they block most escape vectors, but not all.

Attack VectorDockernonoBlocked?
Filesystem escape via bind mount✓-Yes (no bind mounts)
Privilege escalation via setuid✓✓Yes (unprivileged user, no setuid)
Kernel exploit via syscall-✓Partially (syscall filter reduces surface)
Container breakout via Docker socket✓-Yes (socket not mounted)
Network exfiltration via DNS--No (DNS is not HTTP)
Side-channel timing attack--No (not in scope)

Spens does not claim to be a security boundary. It is a “good enough” sandbox for development and testing. If you need defense against a motivated attacker, run Spens inside a VM or use a dedicated sandbox service.

Agent Support

Spens supports four agents out of the box:

  • pi: A coding agent that uses the Anthropic API.
  • opencode: An open-source coding agent.
  • claude: Direct integration with Claude via the Anthropic API.
  • codex: OpenAI Codex integration.

Each agent runs in a container with either a Node.js or Python runtime. You can add custom agents by writing a small adapter that implements the Spens agent interface.

The agent interface is simple:

class Agent:
    def run(self, workspace: Path, prompt: str) -> SessionLog:
        # Execute the agent's logic
        # Return a structured log
        pass

Spens handles container setup, network interception, and log capture. The agent adapter handles LLM API calls and tool invocations.

Session Logs

After a session ends, Spens writes a structured log to disk. The log includes:

  • LLM API calls: request body, response body, token counts, latency.
  • Tool invocations: tool name, arguments, result, exit code.
  • File changes: path, operation (read, write, delete), content diff.
  • Network requests: URL, method, headers, request body, response body, status code.

The web viewer renders this log as a timeline. You can filter by event type, search for specific API calls, and export the log as JSON.

This is useful for debugging agent behavior, auditing autonomous runs, and replaying sessions with different parameters. If an agent makes a mistake, you can see exactly which LLM call or tool invocation caused it.

Deployment Shape

Spens runs on your local machine. It requires Python 3.11+, Docker, and (on macOS) OrbStack. You install it with pip or uv:

pip install spens-ai

You run it from the command line:

spens node-24 pi .

There is no server component, no cloud service, no telemetry. Logs stay on your machine. If you want to share a log, you export it as JSON and send it manually.

This makes Spens easy to adopt but hard to scale. If you need to run hundreds of agent sessions in parallel, you will need to build your own orchestration layer on top of Spens. The tool is designed for local development, not production deployment.

Failure Modes

Spens can fail in several ways:

  • mitmproxy crashes: If mitmproxy crashes, the agent loses network access. The session log will show the crash, but you will not see any network activity after that point.
  • Container OOM: If the agent uses too much memory, the container is killed. Spens logs the OOM event but cannot recover the session.
  • Docker socket unavailable: If Docker is not running or the socket is not accessible, Spens cannot start a container. This is a startup failure, not a runtime failure.
  • Domain allowlist misconfiguration: If you forget to add a required domain to the allowlist, the agent’s API calls will fail. The session log will show 403 errors.

The most common failure mode is domain allowlist misconfiguration. You run an agent, it fails to call an API, and you realize you forgot to add the domain to spens.yaml. The fix is simple, but the error message is not always clear.

Technical Verdict

Use Spens if you are building or testing coding agents and need to see exactly what they do. It is useful for debugging agent behavior, auditing autonomous runs, and understanding which APIs an agent calls.

Avoid Spens if you need a production-grade security boundary, high-throughput orchestration, or multi-tenant isolation. It is a local dev tool, not a platform.

The architecture is straightforward: Docker for isolation, nono for extra sandboxing, mitmproxy for network interception, and a session logger for observability. The trade-off is simplicity over scale. If you need to run one agent at a time and review its behavior afterward, Spens gives you the plumbing.