The context window is the most expensive resource in AI coding. Every token you save is money, speed, and accuracy. The question is how you save them.
Three approaches have emerged in the last year, and they represent genuinely different philosophies about where the savings should come from.
Context Mode — 19.7k★, ELv2
The viral HN #1 darling. An MCP server that sandboxes tool output (98% reduction), persists session state to SQLite, and enforces a “think in code” paradigm — instead of reading 50 files into context, the agent writes a script that computes the result and logs only the output.
The good: The session continuity model is clever — events are indexed into FTS5 and retrieved via BM25, not dumped back into context. The “think in code” enforcement is opinionated but effective.
The catch: ELv2 license. Not OSI-approved open source. Free to use but restricted — you can’t provide it as a service, can’t modify and redistribute freely. For a tool that embeds itself in your agent pipeline, that’s a real consideration.
Lean-CTX — 3.5k★, Apache-2.0, Rust
The newer entrant. A single Rust binary that acts as a context engineering proxy — compresses requests, caches file reads, persists session memory, and provides a live savings dashboard. Zero config, works with 30+ agents.
The good: The single-binary approach is elegant. cargo install lean-ctx and it works. The proxy mode compresses the entire request (system prompt, history, tool results) prompt-cache-safe. The live gain dashboard shows token and USD savings in real time.
The catch: It’s another tool in a space where the problem is already solved by existing stack components. The compression proxy is clever but adds a network hop.
Headroom + RTK — your stack
Headroom MCP compresses tool output before it enters context. RTK prefixes shell commands to strip terminal noise. Both are MIT/Apache-2.0. Already wired into the workflow.
The good: No new install, no new platform, no license concerns. The compression happens at the MCP layer — transparent to the agent. RTK’s shell-level approach means savings on every command without asking.
The catch: More fragmented — two tools instead of one. No session memory persistence (that’s handled by Vestige separately). No pretty savings dashboard out of the box (though tokscale and rtk gain cover this).
The Real Comparison
| Context Mode | Lean-CTX | Headroom + RTK | |
|---|---|---|---|
| License | ELv2 | Apache-2.0 | MIT/Apache |
| Install | MCP plugin | Single binary | MCP + shell hook |
| Compression | 98% (sandbox) | 60-90% (proxy) | 80-90% (MCP + RTK) |
| Session memory | FTS5 SQLite | Built-in | Vestige (separate) |
| Savings dashboard | No | Yes (live) | tokscale + rtk gain |
| Think in code | Enforced | Not enforced | Not enforced |
| Works with | 17 platforms | 30+ agents | Agent-specific |
What I’d Pick
If I were starting from scratch today, I’d probably reach for Lean-CTX. The single-binary, zero-config, Apache-2.0 package is hard to beat. It covers the most ground with the least friction.
But I’m not starting from scratch. My stack already has Headroom compressing tool output, RTK trimming shell noise, and Vestige handling session memory. Adding another tool for marginal gains isn’t engineering — it’s thrashing.
Context Mode is impressive engineering and the viral adoption is deserved. But ELv2 in a tool this core to your pipeline is a real concern. If the license ever changes or the project pivots, you’re stuck.
Lean-CTX is the one to watch. If Headroom ever stops being maintained, that’s my next move. For now, the stack I have works.