You trust your AI agent with your repo. Do you trust it with your SSH keys, your ~/.aws, your dotfiles? Most people running Claude Code, Codex, or OpenCode don’t think about it until it’s too late — and by then the agent has already read everything it could reach.
Greywall (Apache 2.0, Go) is a container-free, deny-by-default sandbox built specifically for AI coding agents on Linux and macOS. No Docker, no VMs — kernel-enforced isolation via Bubblewrap namespaces, Landlock, Seccomp BPF, eBPF monitoring, and a TUN-based network capture.
If you’re evaluating AI coding agents — Claude Code, OpenHands, Codex CLI, whatever — you’re probably doing it wrong. Running them against your own infra is dangerous. Running them manually is slow. Running them unrepeatably is pointless.
Harbor fixes that. It’s a framework from the creators of Terminal-Bench that lets you define sandboxed agent tasks, run evaluations against any agent/model combo, and scale across cloud providers.
Harbor wraps each agent task in an isolated environment (Docker locally, or Daytona/Modal/LangSmith/Novita Sandbox in the cloud). You specify: