The creator built the same production-ready Typeform-style form app twice, once with Claude Code and once with Codex, using an identical Research-Build-Verify prompt, then had Codex itself grade both results anonymously. Codex's scorecard named Claude Code the overall winner, ahead on product judgment (9.4 vs 8.6) and especially pragmatism and efficiency (9.8 vs 5.5, finishing in 5h32m for about $447 versus Codex's 61h47m and about $2,932), while Codex scored higher on architecture and execution (9.6 vs 9.0) and testing and reliability (9.8 vs 8.3). The creator's own hands-on look found Claude Code's app, Formora, was far better in functionality than Codex's app, Rillform, though Rillform had better design. He stops short of a blanket recommendation, noting the result actually contradicts his past tests and that he still uses Codex for roughly 80% of his knowledge-work sessions versus Claude Code for about 20%.
Verdict: Codex, judging the two builds anonymously, ruled Claude Code the overall winner over Codex [16:37]
Key takeaways
Goal: The creator gave Claude Code and Codex the identical slash-goal-command prompt to build a production-ready, originally-branded Typeform alternative [00:29]
Goal: The prompt structured the work into three phases, Research, Build, and Verify [00:27]
Note: Looking back, the creator says he would add a planning phase between Research and Build to map out the whole flow [00:53]
Price: Before either tool was revealed, one build was said to cost about $3,000 over three days, the other about $800 over five hours [00:10-00:13]
+ 30 more takeaways
Price: Claude Code's API billing showed a total cost of about $832 versus Codex costing almost $3,000 [12:46]
Note: The creator later says he believes the $832 figure was a hallucination, with usage stats instead showing roughly $800 [19:10]
Spec: Claude Code used a little over 2 million output tokens while Codex used almost 11.5 million output tokens [12:54]
Spec: Claude Code's cost breakdown totaled 776 million tokens across Fable 5, Opus 4.8, and Opus 5 for about $832 [13:15]
Spec: Claude Code finished in about five and a half hours while Codex took almost 62 hours [13:55]
Spec: Claude Code's run used 35 sub-agents and about 2,800 tool calls, while Codex's run used 126 sub-agents and 32,500 tool calls [14:33]
Reveal: Claude Code built the app called Formora and Codex built the app called Rillform [12:00]
Verdict: Claude Code won on product judgment and scope, 9.4 vs 8.6, making clearer MUST/DEFER decisions while Codex pursued 135 capabilities with less restraint [17:15]
Verdict: Codex won on architecture and execution, 9.6 vs 9.0; Claude Code's contract-first waves produced zero merge conflicts while Codex built a more operationally mature system with immutable revisions, offline recovery, migration safety, concurrency handling, and cloud boundaries [17:41]
Verdict: Codex won testing and reliability by a wide margin, 9.8 vs 8.3, adding cross-browser testing, property tests, and fault injection [18:21]
Spec: Test counts were 296 unit tests, 199 test cases, and 102 browser tests for Claude Code versus 2,339 unit tests, 341 test cases, and 391 browser tests for Codex [18:21]
Verdict: Claude Code won pragmatism and efficiency, 9.8 vs 5.5 [19:02]
Price: Claude Code finished in 5 hours 32 minutes for about $447 using 35 agents, versus Codex's 61 hours 47 minutes and about $2,932 using 126 agents [19:02]
Pro: The creator calls Claude/Fable the 'wise owl,' preferring it for creativity, planning, and brainstorming [15:20]
Pro: The creator finds Codex more obedient, reliably running tests to confirm the job is done [15:20]
For: The creator uses Claude Code for development and planning, and uses Codex, via its Claude Code plugin, for adversarial security review, saying it often finds bugs Claude Code missed [18:01]
For: The creator currently uses Codex to drive about 80% of his knowledge-work sessions versus about 20% for Claude Code, switching as models improve [16:47]
Con: The creator says Codex did not interpret the prompt well or explore creatively enough despite working long and hard, calling it a waste of money [16:05]
Pro: The creator says Formora (Claude Code's build) was far better than Rillform (Codex's build) in functionality, even though Rillform had better design [11:32]
Con: Rillform had a bug where uploading an image for the welcome screen showed 'image preview is unavailable' [03:40]
Con: Rillform's confirm popup appears misplaced on the left side of the screen [04:34]
Con: Rillform's theme studio color and corner radius controls don't appear to change anything when clicked [05:31]
Con: Rillform's webhook, email notification, and partial response integrations appear to be demo UI without full functionality built out [05:46]
Con: Formora had a bug where the account/settings button opens settings but no logout option is visible or reachable [06:59]
Con: Formora had a bug where newly added questions kept showing 'question one' with inconsistent page numbering across views [08:56]
Con: Formora's design/theme editor had no back button to easily navigate out of it [10:03]
Note: The creator says this result contradicts his past tests, where Codex was usually more token-efficient and faster than Claude Code [19:43]
Reason: The clearest gap behind Claude Code's overall win was pragmatism and efficiency, finishing about 11x faster and 6.6x cheaper than Codex [19:02]
Claude Code costed me, if I was using API billing, 832 bucks, which means that Codex costed almost $3,000, which is just insane.
Claude Code finished in five and a half hours for roughly $447, using 35 agents.
How this brief was shaped: Evaluation (review / comparison / unboxing) · confidence Medium
Transcript frames this as a head-to-head comparison from the opening line, same prompts fed to Claude Code and Codex with explicit outcome deltas (3 days vs 5 hours, $3000 vs $800) and a stated goal to reveal what each tool is better at. OCR confirms the shared goal prompt used for both builds, and the 25 percent slice digs into UI/output quality differences, which is comparison evidence rather than a single replicable tutorial.
The lens sets this brief's structure, never its facts — every claim is held to the same citation and fact-check standard.