bench

v2026.09-smoke

claude · haiku · low

77.6overallfull
Post on X

Overall rank

#2 of 2

Cost efficiency rank

#1 of 2

Wall time rank

#1 of 2

Telemetry

full

claude · haiku · low ranks #2 of 2 at 77.6 overall, 11.8 behind codex · gpt-5.6-sol · low. It is also the cheapest scored agent this release, at 1.31 API-equivalent.

Tokens by class

Segment width is dollar share, not token share — where the API-equivalent cost actually went.

Input 1.3K <$0.01Cache read 5.5M $0.55Cache write 236.1K $0.30Output 93.9K $0.47

Tokens in

1.3K

Tokens out

93.9K

Turns

173

Wall time

16m 55s

API-equiv. cost

$1.31

Score / $

59.1

Category scores

UX/UI

(1 tasks)

Frontend

58.3 (1 tasks)

Backend

77.0 (1 tasks)

Planning Audit

79.0 (1 tasks)

Bug Fix

100.0 (1 tasks)

Taste

73.8 (1 tasks)

Signature

57.5 (1 tasks)

Per task

TaskGate passObjectiveSubjectiveScoreRuns
Issue board filters100%
55.6 (55.655.6, n=1)
62.5 (62.562.5, n=1)
58.3 (58.358.3, n=1)
shipshit.dev landing hero100%
57.5 (57.557.5, n=1)
57.5 (57.557.5, n=1)
Codebase audit, planted defects100%
75.0 (75.075.0, n=1)
85.0 (85.085.0, n=1)
79.0 (79.079.0, n=1)
Inventory service100%
100.0 (100.0100.0, n=1)
42.5 (42.542.5, n=1)
77.0 (77.077.0, n=1)
Shipcut pricing page0%
Landing page, five themes0%
88.8 (88.888.8, n=1)
73.8 (73.873.8, n=1)
Fix date range overlap bug100%
100.0 (100.0100.0, n=1)
100.0 (100.0100.0, n=1)
compare vs. grok · grok-4.6compare vs. codex · gpt-5.6-sol · low