If you’re evaluating AI coding agents — Claude Code, OpenHands, Codex CLI, whatever — you’re probably doing it wrong. Running them against your own infra is dangerous. Running them manually is slow. Running them unrepeatably is pointless.
Harbor fixes that. It’s a framework from the creators of Terminal-Bench that lets you define sandboxed agent tasks, run evaluations against any agent/model combo, and scale across cloud providers.
Harbor wraps each agent task in an isolated environment (Docker locally, or Daytona/Modal/LangSmith/Novita Sandbox in the cloud). You specify: