Bring your favorite agents.
Use Claude Code, Codex, Kimi, and GLM with the subscriptions you already have. Run multiple agents together.
The desktop workspace for AI coding agents.
Run Claude Code, Codex, Kimi, and GLM side by side in one desktop app. Keep your subscriptions. Halv cuts wasted context while you work.
Free to try.
Tasks solvedCodex: 7/20
fewer tokens overall
lower cost per correct answer
Measured on the first 20 paired SWE-rebench tasks with Codex. Results vary by task and model.
01 / THE WORKSPACE
Open a project. Choose your agent. Work in chat or native terminals, with your sessions, changes, and savings in one place.




Use Claude Code, Codex, Kimi, and GLM with the subscriptions you already have. Run multiple agents together.
Move between chat and native terminals. Split panes across tasks, organize projects, and return to saved sessions.
Review changes, set approvals, and track agent activity and token savings from your workspace.
We ran the same 20 repository tasks with and without Halv. The repository verifiers accepted 10 Halv answers and 7 vanilla answers.
Token totals include all runs, including failed attempts. Divide by verified answers to see the budget each correct answer took.
| METRIC | Codex | Codex + Halv | HALV CHANGE |
|---|---|---|---|
| Tasks solved | 7 / 20 | 10 / 20 | +3 tasks |
| Total tokens | 63,438,109 | 44,289,991 | -30.2% |
| Recorded cost | $1.84932 | $1.40301 | -24.1% |
| Tokens per correct answer | 9,062,587 | 4,428,999 | -51.1% |
| Cost per correct answer | $0.26419 | $0.14030 | -46.9% |
gpt-5.6-luna · medium reasoning · priority service · Codex 0.152.0
The Halv arm used Crux, RTK, and Halv Engine Headroom together. This benchmark measures the complete toolkit.
An early result from 20 task pairs, published by Halv. Halv used fewer tokens on 12 tasks and more on 8. This is not a guarantee for every run.
Inspect all 40 run records, verifier results, token logs, and SHA-256 checksums. The public verifier checks token totals and task outcomes.
03 / THE HALV ENGINE
Halv helps your agents find the right code, trims noisy command output, and compresses model context as you work.
Duplicated files, stale history, repeated instructions — removed. What changes the answer stays. When a request can't be compressed safely, it passes through untouched.
Callers, references, impact — answered from an index instead of reading the repo a file at a time. Agents ask the index and stop burning context to orient themselves.
Test logs, diffs, dependency trees and install chatter are trimmed to the part that matters before a single token is spent on them.
Halv is built to save you far more than it costs. Get more work from the agent subscriptions you already pay for.
Basic
$2/ MONTH
Every feature, with a monthly savings cap. Start small and see how much wasted context Halv can cut from your daily work.
Unlimited
MOST SUBSCRIBED$10/ MONTH
Every feature, with no cap on savings. Give your agents more room to work, across projects and sessions, all month long.
Track saved tokens in the app before you choose a plan. Your agent subscription price stays the same; Halv helps you get more work from it.
INCLUDED ON BOTH PLANS
FREE TO TRY · CANCEL IN TWO CLICKS · WORKS WITH YOUR CLAUDE, CHATGPT, MOONSHOT OR Z.AI PLAN
Halv is a desktop workspace for AI coding agents. Run Claude Code, Codex, Kimi, and GLM in one app with chat, native terminals, project history, and a live savings meter. Halv combines code navigation, command output filtering, and context compression to reduce wasted tokens while you work.
To make your current allowance last longer, reduce unnecessary context and tool output. Halv combines local context compression, code indexing, and command output filtering in one workspace. Anthropic still controls your account's limits.
Start by reducing unnecessary context: keep instructions focused, give Claude relevant files, and use /clear when starting unrelated work. Halv adds code indexing, command output filtering, and local context compression around Claude Code in a desktop workspace. Use its savings meter to measure the difference in your own work. Savings depend on the task and context.
To make your current Codex allowance last longer, keep requests focused and reduce unnecessary context. Halv helps reduce token waste with code indexing, command output filtering, and context compression. OpenAI controls the actual limits.
Give Codex precise tasks, relevant files, and concise AGENTS.md instructions. Disable MCP servers you do not need. Halv combines code indexing, command output filtering, and context compression to reduce wasted tokens. In Halv's first 20 paired SWE-rebench tasks, Codex with Halv used 30.2% fewer total tokens and 51.1% fewer tokens per correct answer. This small, self-run benchmark does not guarantee the same savings on every task.
Yes. Halv supports Claude Code, Codex, Kimi, and GLM. In Pro mode, each agent runs through its native CLI inside Halv. Sign in with your existing agent account and use the subscriptions you already pay for. API keys are optional for providers that support them.
Halv reduces token usage in three ways. Crux gives agents a code index for navigation and references. RTK filters noisy command output before it enters model context. Halv Engine Headroom compresses model context locally. Together, these tools reduce repeated code exploration and unnecessary context.
In Halv's first 20 paired SWE-rebench tasks, Codex with Halv used 51.1% fewer tokens per correct answer and 30.2% fewer total tokens. Repository verifiers accepted 10 of 20 Halv answers versus 7 of 20 Codex answers. Both arms used the same model, tasks, and verifier. Halv published this small, self-run benchmark and all 40 selected run records.
Savings vary by task, model, and context. The 51.1% result measures total tokens divided by verifier-passing tasks, including tokens from unsuccessful attempts. Halv used fewer tokens on 12 benchmark tasks and more on 8. This Codex benchmark does not establish the same savings for Claude Code or guarantee identical answers.
Halv helps you get more work from your existing agent usage budget. Your provider's fixed subscription price stays the same. Lower token usage can reduce metered API costs, depending on pricing and caching. Compare your measured savings with Halv's price to see whether it pays for your workflow.
Halv combines coding agents, chat, native terminals, project history, approvals, and savings tracking in one desktop app. Crux is the code-index MCP server within Halv's toolkit. With a standalone CLI, you manage that workspace and those tools yourself. Halv brings them together around the agents you already use.
Basic costs US$2 per month with a monthly savings cap. Unlimited costs US$10 per month with no savings cap. Both plans include all features and a 7-day free trial. A card is required at signup. Cancel during the trial to pay nothing; otherwise, your subscription renews automatically.
Halv runs code indexing and context compression on your computer. Your chosen model provider receives the prompts and context needed to answer your requests. Halv also connects to services for account access, billing, updates, and basic app events. Review agent actions and approvals before commands run or files change.
Yes. Halv supports macOS on Apple silicon and Intel, Windows x64, and Linux x64. Linux downloads include AppImage and deb packages. Halv is a desktop app; download and install it on a supported computer to use the workspace.
FROM ZERO TO SAVING
01
macOS, Windows or Linux. No API key required.
02
The CLIs you already use, on the plans you already pay for. Logins stay in the CLI.
03
Savings start on your first prompt, receipt by receipt.
macOS · APPLE SILICON + INTEL
WINDOWS · x64
LINUX · AppImage + deb
SIGN IN AND TRY IT FREE.