The big gap in AI agents isn’t intelligence — it’s hands. An LLM can write a perfect playbook, but if it can’t hold a persistent interactive session, follow a live log, drive the SSH client through a password prompt, or watch the browser as it runs, it’s still stuck translating thought into brittle one-shot tool calls.
I went looking at what’s actually shipping right now to give agents a real computer. Six repos, three distinct layers: full sandboxes, interactive PTY servers, and browser control. Here’s how they stack up.
The Layers
These six tools aren’t competing head-to-head — they occupy three different layers of the stack.
- Full workspace sandboxes:
open-computer-use— a whole Linux machine for the agent. - Persistent interactive terminals:
pty-mcp,PiloTY,tuistory— stateless tool calls become stateful sessions. - Browser control:
open-browser-control,open-browser-use— the agent drives your real browser.
The clever bit is that several of them can be composed. You can bolt a browser controller onto a PTY-backed agent, or drop a browser skill into a sandboxed workspace. But each is strongest alone, and I’ll call that out per tool.
Full Sandboxes
open-computer-use — 114★, Python
Links & Stats 👉 https://github.com/Wide-Moat/open-computer-use
![]()
![]()
![]()
An MCP server that hands an LLM its own Ubuntu box — a per-chat isolated Docker container with bash, Python, Node, Java, a Playwright live CDP browser stream, and 13 built-in skills for documents, charts, and frontend design. It can even spawn a Claude Code sub-agent with its own interactive terminal. Tested to 1,000+ MAU in production, and pluggable into any MCP client: Open WebUI, Claude Desktop, n8n.
My take: This is the “give the agent a whole computer” crowd-pleaser. The per-chat container isolation is the right architecture — nothing leaks between users, and the agent can install anything without wrecking your machine. The transformation warning in the README (dashboard/hosted endpoint going offline during a reorg) is worth noting for anyone planning to rely on it long-term. License is FSL-1.1 (Apache-2.0 future) — source-available, not OSI open source.
Interactive Terminal Servers
pty-mcp — 15★, Go, MIT
Links & Stats 👉 https://github.com/raychao-oao/pty-mcp
![]()
![]()
![]()
An MCP server giving agents real interactive terminal sessions: local shells, SSH, serial ports, and persistent remote sessions that survive disconnects via an ai-tmux daemon. The standout is wait_for — the agent blocks server-side until a regex pattern appears in output, instead of burning cycles with sleep 30 && check_status. Also features settle detection, control keys, bounded-memory ring buffers, and an optional audit log that scrubs credentials.
My take: Tiny but intensely focused. It’s built for sysadmins and network engineers — reboot a server and wait for it to come back, tail logs for ERROR|CRITICAL, drive a router over serial. The audit log with automatic redaction of keys and auth headers is exactly the kind of safety rail a security-conscious agent operator wants. Go means a single static binary. Small stars, but it solves a very real “stop polling, block and wait” problem.
PiloTY — 43★, Python, Apache-2.0
Links & Stats 👉 https://github.com/yiwenlu66/PiloTY
![]()
![]()
![]()
Another MCP PTY server aimed at the same stateless-shell wall. One session = one real interactive terminal that persists across tool calls. Tracks two representations — a raw output stream and a rendered screen/scrollback — and classifies terminal state after each call (running, ready, password, confirm, repl, editor, pager). Ships snapshot tools for cursor-heavy TUIs like vim and top.
My take: The terminal_state classification is the thoughtful touch — it tells the agent not just what printed, but what kind of UI it’s now staring at, so it knows whether to type a command, answer a password prompt, or send an escape key. The README owns the risk honestly: “exposes unrestricted terminal access. Treat it like giving the agent your keyboard.” Apache-2.0 and Python make it the easiest to inspect and extend. More battle-forged than pty-mcp, still young.
tuistory — 348★, TypeScript, MIT
Links & Stats 👉 https://github.com/remorses/tuistory
![]()
![]()
![]()
Positioned as “tmux for AI agents” and “Playwright for TUIs.” Wraps any terminal command (like next dev) in a named background session that agents can read logs from, wait on, type into — and that humans can attach to and detach from at any time, sharing the same terminal state. wait-ready and wait-idle replace the agent’s blind sleep-and-hope.
My take: The most original of the three. It’s not an MCP server — it’s a lightweight CLI layer, so it works with any agent that can run a shell command, not just MCP-aware ones. The human/agent-shared-session model is genuinely useful for dev servers and TUIs. Biggest stars in this category for a reason — clean concept, zero lock-in. (License is declared MIT in package.json but no LICENSE file is committed — a minor hygiene gap to keep in mind if you care about formal licensing.)
Browser Control
open-browser-control — 14★, TypeScript, MIT
Links & Stats 👉 https://github.com/smankoo/open-browser-control
![]()
![]()
![]()
An MCP server that drives your real browser — Chrome or Firefox via an extension — using your existing cookies, sessions, and logins. When the agent hits a sign-in, CAPTCHA, or MFA wall, it drops back to you, then picks up where it left off. Works with Claude Code, Claude Desktop, Cursor, Kiro, any MCP client.
My take: The “your real browser, your real logins” approach means the agent can actually do logged-in tasks you’d otherwise have to hand it. The human-in-the-loop handoff for CAPTCHA/MFA is the pragmatic answer to the one thing automated browsing fundamentally can’t do. Both Chrome and Firefox from one protocol is nice for Chrome-bloat fatigue. Young (14★) but the model is sane.
open-browser-use — 231★, JavaScript, MIT
Links & Stats 👉 https://github.com/iFurySt/open-browser-use
![]()
![]()
![]()
A platform-neutral browser automation layer that pairs a browser extension with a CLI, exposed through JavaScript, Python, and Go SDKs. Positioned explicitly as an open-source alternative to the Chrome Browser Use capability now in Codex.app — no agent-runtime lock-in.
My take: The anti-lock-in play, and the one with the most deliberately polyglot surface (JS + Python + Go SDKs, plus CLI). It’s a layer, not a full agent, which makes it the most reusable building block if you already have an agent stack you like. 231★ is the healthiest adoption of the browser pair. Homepage blog post “Browser Use Deep Dive” is worth a read for the backstory.
The Two Pickpockets in the Room
Two things stand out across all six:
- Polling is dying. Every terminal tool here ships a
wait_for/wait-ready/wait_outputprimitive so agents block-and-wait server-side instead ofsleep-looping. That’s the single biggest quality-of-life (and cost-of-API-calls) improvement in this space right now. - Stateless-session pain is the real driver. They all exist because agent tool calls start fresh — no env, no cwd, no live process. The ones that actually solve it (PiloTY’s terminal_state, pty-mcp’s persistent sessions, tuistory’s shared sessions) are the ones worth your time.
How They Stack Up
- Give an agent a whole sandboxed machine →
open-computer-use - Persistent local/SSH/serial terminal with wait semantics →
pty-mcp - Stateful terminal with UI-aware classification, most hackable →
PiloTY - Zero-lock-in background TUIs shared with humans →
tuistory - Drive your real logged-in browser →
open-browser-control - Platform-neutral browser building block →
open-browser-use
For my own stack — headless VPSes, no GUI, terminal-native, event-driven lean — the terminal trio is where the action is. pty-mcp’s audit log and wait-for-regex fit my “no telemetry, no polling, security-conscious” rules best. tuistory is the one I’d reach for in any dev-loop where a human and agent share a process. The browser pair is more useful on a machine with a display, but the handoff-to-human MFA pattern in open-browser-control is the right idea for when logged-in web work can’t be avoided.