bench

latest release

v2026.09-smoke

v2026.09-smoke · suite v2026.09 · generated Sep 6, 2026 · methodology

Overall score

codex · gpt-5.6-sol · low89.5claude · haiku · low77.6

Overall score — higher is better.

API-equivalent cost

claude · haiku · low$1.31codex · gpt-5.6-sol · low$18.45

API-equivalent cost per full run — lower is better.

Wall time

claude · haiku · low16m 55scodex · gpt-5.6-sol · low28m 31s

Wall time per full run — lower is better.

codex · gpt-5.6-sol · low leads at 89.5, claude · haiku · low is 14x cheaper at 77.6; grok · grok-4.6 did not run.

Score vs. cost

API-equivalent cost per full run set against overall score, log scale. The shaded corner is cheap and good; the dashed line is the Pareto frontier — agents no other agent beats on both axes at once.

0255075100$0.83$2.72$8.91$29.25scoreAPI-equivalent cost (log)claude · haiku · lowcodex · gpt-5.6-sol · low
claudecodexPareto frontier

Changelog

  • Sep 6, 2026v2026.09-smoke published — 3 agents, 7 tasks.
  • Sep 6, 2026Task added: Inventory service (backend).
  • Sep 6, 2026Task added: Fix date range overlap bug (bugfix).
  • Sep 6, 2026Task added: Issue board filters (frontend).
  • Sep 6, 2026Task added: Codebase audit, planted defects (planning-audit).
  • Sep 6, 2026Task added: shipshit.dev landing hero (signature).
  • Sep 6, 2026Task added: Landing page, five themes (taste).
  • Sep 6, 2026Task added: Shipcut pricing page (ux-ui).