The fx coding agent, Vercel's open-source Unix-like terminal agent, is getting faster and more memory-efficient. According to the project's public README on GitHub, fx now supports an Ultrafast mode for latency-sensitive requests, alongside smarter session compaction that keeps recent tool exchanges intact when long conversations overflow the context window.

The update lands as coding agents settle into a second, more demanding phase of adoption. The first wave was about getting agents to work at all; this one is about making them fast, cheap, and reliable enough to run all day inside real development workflows. For teams that live in the terminal, a mode that buys speed and a compaction system that loses less state both matter more than another model bump.

What is the fx coding agent?

fx is an open-source, native coding agent built by Vercel's labs team. According to the project's documentation, it is written in the Zig programming language and designed to feel like a Unix tool: a single native binary that runs in the terminal, speaks the Agent Client Protocol (ACP), and can also be embedded as a library, called libfx, in Node.js, browsers, and Next.js apps.

The agent is deliberately provider-agnostic. According to the README, fx sessions can switch between the Vercel AI Gateway, Codex subscription access, and Grok subscription access with a single command, and users can wire up custom model connections — including local servers such as Ollama or gateways such as OpenRouter — through a settings file. The repository's public page shows it has gathered thousands of stars since launch, according to GitHub.

fx also tracks its own costs. According to the documentation, a built-in usage command reads locally recorded token spend, while Gateway users can check credit balances from inside a session — small touches that reflect how agent developers actually budget their work.

Ultrafast mode, explained

The headline addition is Ultrafast mode, which routes requests through OpenAI's higher-cost Gateway service tier. According to the project's README, the mode is off by default and requests the Ultra tier for models whose Gateway metadata advertises Ultra eligibility. A status command reports whether Ultra was requested — carefully noting that this is a request, not a guarantee that the provider served it.

Ultrafast can be turned on per session, per profile, or per command. According to the documentation, users can set a default in their settings file, pass a command-line flag for one-shot requests, or toggle it interactively inside the shell. Notably, background side calls — including titles, reviews, and compaction itself — do not use Ultra mode, according to the README, which keeps the expensive tier reserved for the turns where latency actually matters.

The design reflects an honest tradeoff. Speed at the provider level costs more, so fx makes it opt-in and scoped: fast turns pay the premium, background housekeeping does not. The fx coding agent thus gives developers a deliberate lever over latency rather than a blanket upgrade.

Compaction gets smarter

The second upgrade is quieter but arguably more important for day-long sessions. When a conversation approaches the context limit, fx automatically compacts history — summarizing the older part of a session so the agent can keep working in a fresh window. According to the v0.0.8 release notes, compaction now keeps recent tool exchanges intact, preserves the full transcript, and continues the same turn rather than restarting it.

According to the v0.0.9 release notes, compaction is now visible in the activity row of the terminal UI, and sending a message while it runs no longer races it — the message waits, then runs against the freshly compacted context. The same release fixed an edge case where interrupted work or failed file lookups could leave subagents stuck mid-run.

A deeper technical write-up of the fx backend by the agent-sdk project describes the compaction algorithm in detail: it walks backward through history to the most recent user messages, generates a summary of everything older — typically about ten percent of the input length — then rewrites the session log atomically so a crash mid-compaction never corrupts the record. According to that documentation, todos are re-injected into the compacted history so nothing about the agent's plan gets lost.

Why it matters for agent developers

fx's trajectory says something about where coding agents are headed: toward native, inspectable tooling that treats context and cost as first-class engineering problems. Ultrafast mode acknowledges that some turns — an interactive review, a live demo — are worth paying more for, while compaction work acknowledges that long-running agents live or die on memory management.

The Unix philosophy shows up everywhere. According to a write-up by Script by AI, fx treats sessions, providers, and permissions as composable command-line operations: resume a conversation, compact on demand, or embed the whole agent as a library. Agents-as-Unix-tools is a compelling model precisely because developers already know how to script, pipe, and automate the terminal.

It also reflects the broader consolidation of model providers behind gateways and ACP-compatible clients. As more teams standardize on agent protocols, the differentiator shifts from raw model quality to the quality of the harness: how fast it responds, how gracefully it handles hundred-turn sessions, how clearly it reports what it spent. As Vercel's own repository demonstrates, that harness is now being built in the open.

Related coverage on this site includes Reflection AI's open-weight model bid and New York's AI safety hearing, which continue to shape the competitive landscape that agents like fx run in.