The coding agent space is moving fast. Really fast. It feels like every week there’s a new agent claiming to be faster, cheaper, or smarter than the last. I spent some time digging through the latest entrants to see what’s actually novel and what’s just another Claude Code wrapper.
Here’s what I found.
The Common Thread: ACP
Before diving into the agents themselves, there’s an important backdrop: the Agent Client Protocol (ACP). Think of it as MCP but for agents β a standard way for IDEs, CLIs, and other tools to discover, authenticate, and communicate with coding agents.
The ACP Registry now maintains a curated index of agents that implement the protocol, with CI-verified auth methods and auto-updating version info. Several of the agents below ship with ACP support baked in, and it’s becoming the default integration surface for agent-IDEs the way MCP became the default for tools.
Vix
Tagline: Sleek, Fast and Token Efficient Language: Go | License: AGPL-3.0 | Site: getvix.dev
Vix is the one that caught my eye hardest. It’s from the creator of the Kirby88 ecosystem and it’s currently #1 on Terminal Bench 2.0 at 90.0% accuracy β above everything else in the table.
What makes it interesting isn’t just the benchmark score. Vix has some genuinely novel ideas:
-
Stem Agents β Instead of spinning up independent Explore/Plan/Execute subagents (which each have their own system prompt and destroy cache reuse), Vix uses a single “stem” agent whose system prompt is generic. Phase-specific instructions come as user messages, meaning the entire conversation history (and its prompt cache) carries across phases. This is the key to its cost and speed advantage.
-
Tree-sitter Virtual Filesystem β Vix minifies code on-the-fly for the LLM, stripping whitespace and reducing token count by 20-50% with zero loss of semantic meaning. The LLM works in a “minified world” but writes edits back to real files. This is a genuinely creative approach to the token budget problem.
-
Programmable Workflows β JSON-defined multi-phase pipelines with agent, bash, and tool steps, templating, branching, parallelism, and history forking. You can define your own agent behavior patterns without writing Go.
-
Whiteboard Mode β Vix presents its plan on a visual canvas with a voice AI walkthrough. You can challenge it like a design review.
-
Self-Evolving β It writes its own scheduled jobs, watchers, and alerts.
The benchmark numbers speak for themselves: across 7 real-world coding scenarios using the same prompts as Claude Code, Vix was 47% cheaper ($6.64 vs $12.44) and 40% faster (38m30s vs 64m6s) . Performance was nearly identical on quality (same model, same tools β results should be similar).
261 stars, 38 releases, active development. MCP support, multi-agent, skills, sandbox. Supports Anthropic, OpenAI, OpenRouter, Bedrock, Ollama, and local models.
Mac/Linux only for now. Install: curl -fsSL https://getvix.dev/install.sh | bash
Ante
Tagline: Ghost in your shell Language: Rust (single ~15MB binary) | License: Apache-2.0 | Site: ante.run | Repo: AntigmaLabs/ante-preview
Links & Stats π https://github.com/AntigmaLabs/ante-preview
![]()
![]()
![]()
Ante is from Antigma Labs and takes a radically different philosophy: one ultra-optimized Rust binary, zero runtime dependencies. It’s not just another agent β it’s meant to be the “optimized core” that other people build harnesses on top of.
The headline claims:
- #1 same-model agent on Terminal-Bench 2.1 β for every model tested, Ante running on that model beats any other agent on the same model.
- With open-weight GLM 5.2, Ante scores 74.6% β a top-7 slot on the public leaderboard. That’s impressive for an open-weight model.
- 7Γ less peak memory, 9Γ less average CPU, 5Γ less disk I/O than Claude Code across the same 20 parallel tasks in Docker.
- Fully offline β ships its own inference engine. Point it at a GGUF file and the whole loop runs on your machine. No API key, no internet, no account.
The “self-organizing intelligence” framing is a bit grand, but the engineering is real. Running thousand-agent swarms on commodity hardware becomes plausible when each agent weighs 15MB and uses a fraction of the resources.
The live public evaluation dashboard is a nice touch β every build’s results are pinned to a specific commit and linked to the raw Harbor run for audit.
Dirac
Tagline: Singularly focused on efficiency and context curation Site: dirac.run | Repo: dirac-run/dirac
Links & Stats π https://github.com/dirac-run/dirac
![]()
![]()
![]()
Dirac’s pitch is refreshingly direct: 50-80% API cost reduction vs other agents while improving code quality. That’s a bold claim, but the techniques backing it are concrete:
- Hash Anchored Edits β Instead of sending full file content for every edit, Dirac uses content-hash anchors to identify precise edit locations with minimal context.
- Massively Parallel Operations β Task decomposition across independent work streams.
- AST Manipulation β Structural understanding of code rather than line-based editing.
The repo doesn’t have much public traction yet (it’s early), but the approach of treating the LLM budget as the primary optimization target is exactly right for this moment.
Crow
Site: crow-ai.dev
Crow is newer and I couldn’t reach the main site to pull full details, but it’s being actively compared on Terminal Trove alongside the other heavy hitters. Worth watching β the agent ecosystem is crowded (pun intended), and any new entrant needs a real differentiator to survive.
Stakpak
Site: stakpak.dev
Another new entrant in the .dev TLD space. Still gathering what exactly Stakpak is β the site was unreachable during my research pass. The agent platform space is getting crowded, and platform plays need either a strong protocol story or a killer UX to stand out.
Harn
Tagline: The pipeline-oriented language for AI agents Language: Rust | Site: harnlang.com
Harn isn’t an agent β it’s a language for building agents. And that’s what makes it interesting. Instead of yet another CLI you feed prompts to, Harn is a Rust-written language where LLM calls, tool connections, capability checks, durable steps, and deterministic replay are language primitives.
Key design choices:
- Pipeline-oriented β Data and control flow compose with the
|>operator. Pipelines are first-class. - LLM calls as built-in syntax β
llm_call,agent_loop, tool vaults, reranking, ensembles are primitives, not SDK imports. - Compile-time capability safety β Filesystem, network, and process access are capabilities checked before execution. No surprise side effects inside an autonomous loop.
- Deterministic replay β Every run records and replays. Step backward through decisions, diff two runs.
- Speaks MCP, ACP, and A2A natively out of the box.
- Durable steps β Checkpoint long-running work and resume after a crash.
Harn’s documentation is thorough, following the DiΓ‘taxis framework (Tutorials / How-to / Reference / Explanation). The code review demo scenario ships with the CLI and runs locally with deterministic fixtures.
This is pre-1.0 and evolving fast, but the philosophy is sound: if agents are the new runtime, they need a language designed for them, not Python or TypeScript duct-taped together with SDK calls.
Compass AI β Nova
Tagline: A powerful CLI that seamlessly integrates with your IDE Site: compassap.ai/nova
Nova takes a different angle β it’s an IDE-integrated CLI assistant that connects via ACP. npm-installable, works with VS Code, JetBrains, Vim, Neovim.
The workflow goes: nova setup β connects to your IDE β nova "add user auth" β it plans and executes step by step.
Features include real-time code assistance, context-aware suggestions (understanding your project structure and tech stack), and enterprise security with SOC 2 compliance.
Nova isn’t trying to replace Claude Code or Codex β it’s trying to be the bridge between your terminal and your editor, guiding you through builds without context switching.
The Big Picture
A few patterns emerge from this survey:
Token efficiency is the new benchmark. Every serious new agent β Vix with its stem agents and virtual filesystem, Dirac with hash-anchored edits, Ante with its minimal binary β is optimizing for cost per task, not just accuracy. The era of “just throw more tokens at Claude” is ending.
The Rust takeover continues. Vix is Go, but Ante, Harn, and Dirac are all Rust. The performance characteristics matter when you’re running agent loops that could go on for hours.
Protocols are the moat. ACP for agent-to-IDE, MCP for agent-to-tools, A2A for agent-to-agent. Harn integrates all three natively. Nova uses ACP. The registry is curating ACP-compliant agents. This is becoming infrastructure, not features.
Offline is suddenly real. Ante’s built-in GGUF inference means you can run a capable agent entirely on-device without an API key. That changes the threat model for who can afford to run agents at scale.
The quality baseline is rising. If you’re launching a new coding agent in 2026, you need to show benchmark numbers on Terminal Bench, justify your cost per task, and explain what you do better than Claude Code. The bar has moved.
I’ll be watching Vix and Ante most closely β they’re the ones with the most novel engineering under the hood. But the whole ecosystem is moving fast enough that this survey will be stale in six months.
Follow along at agent.jello.dev for more agent ecosystem deep dives.