← Back to Pulse PULSE. brief
I Tested Claude Code vs. Codex on Design. It Wasn't Even Close.

I Tested Claude Code vs. Codex on Design. It Wasn't Even Close.

Nate Herk | AI Automation20 min2026-08-26 ▶ Watch on YouTube
What this video is
⚡ a 21-minute video, readable in 60 seconds

This is a build-off between Claude Code and Codex across eight matched website builds, judged on design plus sub-agent count, time, cost, and output tokens rather than vibes. The reviewer reaches no single overall design winner and calls it mixed: Codex takes the first build outright, leads two wins to one after three rounds, and produces near-identical output to Claude Code once prompts get very specific, while Claude Code gets the nod on parts of the TrailLatch build. On efficiency the reviewer is not mixed at all, repeatedly calling Codex faster, cheaper, and more token-efficient, and the final tally puts total API cost at about $444 for Claude Code versus about $100 for Codex.

Verdict: the creator sets out to judge which AI produces better designs across eight matched builds, scored on sub-agents, time, cost, and output tokens rather than vibes [00:11]
Key takeaways
+ 43 more takeaways
  • Con: the Codex-built Bowl & Bloom cart let users adjust quantities but did not update the matching color-coded display [02:52]
  • Vs: reviewer names Codex the clear winner on design, feel, cost, and time for the Bowl & Bloom round [03:35]
  • Con: reviewer calls Claude Code's MinutesCraft pricing page 'over building,' more overwhelming with seats, extra seats, and solo practice options [06:12]
  • Pro: the Codex-built MinutesCraft version shows two clear action items, a draft supplier brief for Taylor and a pilot score card shared with Mina [05:42]
  • Con: reviewer spots an obvious UI image bug on the Claude Code MinutesCraft tour page in full screen mode [06:30]
  • Spec: Codex took 1 hour 12 minutes and used 100,000 tokens costing $20 on the MinutesCraft build [07:31]
  • Spec: Claude Code took 2.5 hours and used almost half a million tokens costing $65 on the MinutesCraft build [07:31]
  • Vs: reviewer reacts to the Claude Code-built TrailLatch design saying 'Okay, honestly, I really like this' [09:10]
  • Vs: reviewer reacts to the Codex-built TrailLatch design saying 'Right off the bat, I do not like this,' later calling it hard to read and confusing [09:16]
  • Con: reviewer calls the TrailLatch 'Choose the Trip' page overwhelming, saying there is too much going on [10:22]
  • Spec: on the TrailLatch build, Claude Code used four sub-agents versus one for Codex [11:11]
  • Spec: Claude Code used almost 400,000 tokens and cost $42 on the TrailLatch build [11:17]
  • Spec: Codex used about 100,000 tokens and cost a little under $20 on the TrailLatch build [11:17]
  • Verdict: reviewer states Codex is faster, more efficient, and cheaper than Claude Code [11:26]
  • Vs: reviewer scores the design comparison two wins to one in favor of Codex after three builds [11:35]
  • Vs: reviewer says that over a longer back-and-forth session, Codex would likely follow instructions more efficiently than Claude Code [12:08]
  • Vs: reviewer predicts 'I think Codex is going to win this one for sure' going into the Present & Clear build [13:41]
  • Spec: the four-agent Claude Code approach took 2 hours 9 minutes, used 442,000 output tokens, and cost $42 on the Present & Clear build [13:45]
  • Spec: the single-agent Codex approach took 48 minutes, used 92,000 output tokens, and cost $18.80 on the Present & Clear build [13:45]
  • Pro: reviewer says he is liking Codex better so far, citing buttons, a logo, and a more branded, trusted feel [14:10]
  • Con: reviewer describes the Claude Code version as having a lot of words and not very many visuals [14:20]
  • Pro: the Codex version has fewer words and more structure [14:23]
  • Pro: reviewer calls a scrolling path element on the Codex version 'cool' [14:27]
  • Con: reviewer reacts to a North Ledger Studio version saying 'this just feels like a report' and questions who would want to look through it [14:47]
  • Pro: reviewer praises the branding of a North Ledger Studio version as 'a much better branded feel' than the other [15:00]
  • Pro: reviewer calls the Claude Code-built Roomtone hero section 'not a bad hero section' though a little overwhelming [15:45]
  • Con: reviewer criticizes the Claude Code-built Roomtone site as too wordy, with room dimension numbers that looked visually stretched [16:28]
  • Spec: on this open-ended build, Codex finished in 8 minutes for about $1.50 [17:41]
  • Spec: Claude Code took nearly 3 hours and 330,000 tokens for about $50 on the same open-ended build [17:41]
  • Vs: on a highly specific Basalt prompt, reviewer says the Claude Code and Codex outputs are 'pretty much the exact same,' with only small differences like slightly different icons [18:15]
  • Verdict: reviewer says being more specific in a prompt lets you get basically any model or harness to produce exactly what you're looking for [18:24]
  • Spec: Claude Code took 22 minutes, used 110,000 tokens, and cost $14 on the Basalt build [18:48]
  • Spec: Codex took 6 minutes, used 16,000 tokens, and cost $1.08 on the Basalt build [18:48]
  • Spec: Claude Code was run on Opus 5 while Codex was run on GPT 5.6 [18:53]
  • Vs: on the incident postmortem report, reviewer again concludes the two outputs are 'pretty much the exact same' from a design perspective [19:06]
  • Verdict: reviewer says the cost was significantly better for Codex despite similar output quality on the incident report [19:20]
  • Spec: totaled across all eight builds, Claude Code used 25 sub-agents versus 9 for Codex [19:51]
  • Spec: Claude Code ran for 14 hours total versus 5 hours for Codex [19:56]
  • Spec: Claude Code used almost three million output tokens across all builds [19:59]
  • Spec: Codex used about 550,000 output tokens across all builds [20:04]
  • Price: at API billing rates, Claude Code would have cost about $444 in total across all builds [20:07]
  • Price: at API billing rates, Codex would have cost about $100 in total across all builds [20:07]
  • Bowl & Bloom shows box size selection (8/$88, 12/$126, 16/$160) paired with mix customization, displaying ingredient details and a summary panel with color-coded meal counts.
How this brief was shaped: Evaluation (review / comparison / unboxing) · confidence Medium

Transcript shows one creator running a head to head comparison, having Claude Code and Codex build eight identical websites from the same brand prompt and judging which is better on design while backing it with sub agent counts, build time, cost, and output tokens rather than just vibes. This is a comparison with a verdict, matching evaluation, and confidence is capped because the OCR sample is about an unrelated freezer meal subscription box and does not match the transcript, so it cannot confirm the on-screen site comparisons.

The lens sets this brief's structure, never its facts — every claim is held to the same citation and fact-check standard.

Jump to a moment
Their links, sorted & clickable
🏛️ Communities & courses1My FREE resourcesskool.com
🤝 Work with them1My playbook for growing a $1M AI agencyapp.aiautomationsociety.ai
🛠️ Tools they use1FREE MONTH voice to textget.glaido.com
📢 Sponsored / affiliate1Code NATEHERK for 10% off VPS (annual plan)hostinger.com
💼 Sponsorship & business1Sponsorship Inquiries
🌐 Find them3LinkedInlinkedin.comX / Twitterx.comInstagraminstagram.com
🔗 Other links1Try Granola for FREE todaygranola.ai
← Back to Pulse Dashboard
Was this brief useful?