bench
shipshit.dev model benchmark
Leaderboard
Tests
Methodology
Source
v2026.09-smoke
· compare
claude · haiku · low
vs.
codex · gpt-5.6-sol · low
Post on X
claude · haiku · low
77.6
$1.31 API-equiv.
full
codex · gpt-5.6-sol · low
89.5
$18.45 API-equiv.
partial
Every task
Issue board filters
58.3
(58.3–58.3, n=1)
run
68.3
(68.3–68.3, n=1)
run
shipshit.dev landing hero
57.5
(57.5–57.5, n=1)
run
92.5
(92.5–92.5, n=1)
run
Codebase audit, planted defects
no screenshot
79.0
(79.0–79.0, n=1)
run
no screenshot
79.0
(79.0–79.0, n=1)
run
Inventory service
no screenshot
77.0
(77.0–77.0, n=1)
run
no screenshot
—
run
Shipcut pricing page
no screenshot
—
run
100.0
(100.0–100.0, n=1)
run
Landing page, five themes
73.8
(73.8–73.8, n=1)
run
100.0
(100.0–100.0, n=1)
run
Fix date range overlap bug
no screenshot
100.0
(100.0–100.0, n=1)
run
no screenshot
100.0
(100.0–100.0, n=1)
run