Halv used 51% fewer tokens per correct answer on SWE-rebench
On 20 paired SWE-rebench tasks, Halv used fewer tokens, cost less, and solved more tasks than vanilla Codex under the same model and verifier.
Read postHALV BLOG
News about Halv and notes from the team building it.
On 20 paired SWE-rebench tasks, Halv used fewer tokens, cost less, and solved more tasks than vanilla Codex under the same model and verifier.
Read postThirteen tested ways to cut token usage in ChatGPT and Codex: the credit rate card, model and effort choice, fast mode, caching, and AGENTS.md.
Read postFourteen tested ways to cut token usage in Claude Code and claude.ai: context hygiene, model and effort choice, MCP pruning, hooks, and caching.
Read postHow Codex usage limits work in 2026: the five-hour window that came back for Plus on August 25, the weekly cap every plan keeps, why Pro is exempt, what drains your allowance fastest, and how to get more from your ChatGPT plan.
Read postHow Claude usage limits work in 2026: the 5-hour session window, weekly caps, what burns tokens fastest in Claude Code, and how to get more from your plan.
Read postThe follow-up benchmark: Crux in real multi-question coding sessions on django and sympy. 90% correct answers versus grep's 58% — at 47% less cost per correct answer.
Read postMethodology and full data for the session-shaped Crux benchmark, the 2.3x token blowup it exposed on sympy, and the Crux 0.6.2 output redesign that turned it into a 59% saving.
Read postWe benchmarked an AI coding agent with and without Crux, our code-index MCP server, in Codex on django and sympy: 45% more correct answers, 24% less cost per answer.
Read postMethodology and full data for the Crux benchmark: an SCIP code-index MCP server vs grep in Codex CLI, with token-metered paired runs, ablations, and the tuning trail.
Read post