mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

Dev Tools

Token Compression for Coding Agents: Fine-Tuned Middleware Cuts Codex Costs by 30%

A Show HN project uses a fine-tuned Qwen model as a proxy layer to compress tool-call output, reducing input tokens and API spend for coding agents.

Source: news.ycombinator.com
Token Compression for Coding Agents: Fine-Tuned Middleware Cuts Codex Costs by 30%

Coding agents hit a cost wall when tool-call output bloats context windows. A Show HN project tackles this with a fine-tuned compression model that sits between agent output and model input, trimming tokens by 29.6% without breaking KV cache or multi-turn reasoning. The project exposes a pattern: developers are building custom middleware layers to manage the economic pressure points in agentic workflows.

The Problem: $700/Day API Bills

The team behind this project maxed out their Codex subscription and burned $700 per day per person on API calls. The culprit was not the agent’s reasoning steps but the tool-call results that get fed back into the model. File retrieval, test output, and error traces accumulate fast. Each round trip inflates the input token count, and cache misses compound the cost.

Coding agents differ from conversational agents in context shape. A chat agent might reference a few messages. A coding agent drags in file trees, diff output, stack traces, and test results. The context window fills with structured data that the model needs for the next step but that also contains redundancy.

Architecture: Proxy Layer with Fine-Tuned Compression

The solution is a local proxy that wraps Codex and intercepts tool-call results before they return to the model. The proxy runs a fine-tuned Qwen model trained to preserve agent trajectory while removing redundant information from tool outputs.

Execution flow:

  1. Agent calls a tool (file read, test run, search).
  2. Tool returns structured output (file content, test logs, search results).
  3. Proxy intercepts the output and passes it to the compression model.
  4. Compression model trims redundant tokens while preserving semantic fidelity.
  5. Compressed output goes back to the agent’s context window.
  6. KV cache remains untouched because the proxy operates before the model sees the input.

The fine-tuning objective is trajectory preservation. The model learns which parts of tool output the agent needs for subsequent reasoning steps and which parts are noise. File retrieval accuracy and context-heavy tasks see the biggest gains.

Implementation Details

The CLI is open source and installs as a shell wrapper around Codex. It runs on by default, and you can disable it with codex --uncompress when you need full output.

Key design choices:

  • Local proxy: No data leaves your machine. The proxy wraps your local Codex instance and does not retain queries.
  • Fine-tuned Qwen model: The compression model is trained on coding agent trajectories, not general text. This preserves the structure that agents need for multi-turn reasoning.
  • KV cache preservation: The proxy compresses tool output before it enters the model’s context window, so the cache does not invalidate.
  • Token counting: The proxy uses OpenAI’s response.usage to measure savings. You can run savings to see cumulative token reduction.

Installation:

curl -fsSL https://install.everestagi.com/install.sh | sh && \
  source ~/.config/everest/shell.sh

The proxy runs as a local service and intercepts API calls. You point your Codex client at the proxy endpoint instead of the OpenAI endpoint.

Trade-Offs and Failure Modes

DimensionBenefitRisk
Cost29.6% token reduction, lower API spendCompression model adds latency and local compute overhead
FidelityFine-tuned on agent trajectories to preserve reasoning stepsMay remove information the agent needs for edge cases or complex tasks
CacheOperates before model input, so KV cache stays validIf compression changes output shape, downstream tools may break
PrivacyLocal proxy, no data retentionRequires trust in the proxy binary and shell script installer
PortabilityWorks with any Codex-compatible clientTied to Codex and Astra workflows, not general-purpose

Failure modes:

  • Over-compression: The model removes a file path or error message that the agent needs for the next step. The agent halts or makes incorrect assumptions.
  • Cache invalidation: If the compression model changes output structure in a way that breaks tool-call contracts, the agent’s cache becomes stale.
  • Latency: Running a fine-tuned model locally adds milliseconds to each tool call. For high-frequency agents, this accumulates.
  • Model drift: If Codex changes its tool-call format, the compression model may need retraining.

Observability and Debugging

The proxy exposes a savings command that shows cumulative token reduction. This is useful for tracking ROI, but it does not expose per-call compression ratios or failure cases.

What you cannot see:

  • Which tool calls benefit most from compression.
  • When the compression model removes information that causes downstream errors.
  • Latency breakdown (tool call, compression, model inference).

For production use, you would want structured logs that capture input tokens, output tokens, compression ratio, and agent success rate per task type. You would also want a fallback mode that disables compression if the agent fails repeatedly on a specific task.

When to Use This

Good fit:

  • You run coding agents daily and hit API cost ceilings.
  • Your tasks are context-heavy (file retrieval, test output, large diffs).
  • You can tolerate local compute overhead and occasional compression errors.
  • You trust the proxy binary and are comfortable with shell script installers.

Poor fit:

  • You run conversational agents with small context windows.
  • Your tasks require exact tool-call output (legal, compliance, security).
  • You need sub-100ms latency and cannot afford local model inference.
  • You operate in environments where local proxies violate security policy.

Technical Verdict

This project demonstrates a practical response to the economic pressure points in agentic workflows. The architecture is sound: a local proxy with a fine-tuned compression model preserves KV cache and reduces input tokens without breaking multi-turn reasoning. The 29.6% token reduction is meaningful for teams burning hundreds of dollars per day on API calls.

The trade-off is fidelity. Fine-tuning helps, but compression always risks removing information the agent needs. For production use, you would want observability that tracks compression ratio per task type and a fallback mode that disables compression when the agent fails.

Use this if you are optimizing for cost and can tolerate occasional compression errors. Avoid it if you need exact tool-call output or operate in environments where local proxies are not allowed. The pattern is worth watching: as agents move from prototypes to daily workflows, cost-optimization middleware will become a standard layer in the stack.