A head-to-head test of Claude Code and Codex building eight websites from identical briefs, comparing design quality, agent count, build time, tokens, and API cost.
Adapted from @nateherk# I Tested Claude Code vs. Codex on Design. It Wasn't Even Close. I gave Claude Code and Codex the same prompts, brand guidelines, copy, logos, and source data to build eight websites. Then I reviewed every result side by side and tracked the agents, time, output tokens, and estimated API cost behind each build. Codex won most of the subjective design rounds for me, but the way I wrote the prompt changed how large the gap was. TL;DR → Codex won five of the six open-ended design rounds. Claude Code won one. → Codex used nine agents and finished in about five hours. Claude Code used 25 agents and took about 14 hours. → Codex used roughly 550,000 output tokens. Claude Code used almost three million. → Equivalent API billing came to about $100 for Codex and $444 for Claude Code. → When I specified nearly every design decision, the outputs looked almost identical. ## The Setup I wanted to remove as much variability as possible. For the first five builds, both tools received the same business concept, copy, logo, brand guidelines, and website design skill. Claude Code ran Opus 5. Codex ran GPT-5.6 Sol. Responsive behavior was part of the test because a website can look great at full width and fall apart when the window gets smaller. I judged the experience, but I also recorded what happened under the hood: → Number of agents → Elapsed build time → Output tokens → Estimated API cost Design taste is subjective. Time, tokens, and cost give the comparison another layer. ## Five Head-to-Head Landing Pages ## 1. Bowl and Bloom Bowl and Bloom was a postpartum meal-delivery business. Both versions used the same colors, story, four core flavors, and care-box concept, but they communicated the offer very differently. Claude Code opened with a lot of text and turned the page into a chapter-based story. Codex put the product image first, added clear navigation and calls to action, and used a progress indicator with a horizontal-scroll section. The Codex version made it easier to understand the offer as soon as the page loaded. Both tools built a care-box configurator. Claude Code had a layout issue in half-screen view. Codex handled the smaller viewport better, although its selected quantities did not update the visible bowl counts correctly. Codex still won the round for me. → Claude Code: four subagents, about two hours, almost 500,000 output tokens, about $53 → Codex: one agent, about 50 minutes, about 100,000 output tokens, just under $20 ## 2. MinuteCraft MinuteCraft was an AI meeting notetaker that turns calls into decisions and action items. Both versions felt more like software dashboards than landing pages built to convert. Claude Code made the problem worse by adding more copy, more steps, and more technical-looking sections. Codex organized the sources, summary, transcript, notes, and assigned action items more clearly. The gap became larger on the secondary pages. Claude Code's tour layout was broken even at full screen. Its pricing and integrations pages were dense enough that I did not know where to look first. Codex used fewer words, cleaner columns, and a simpler trial flow. → Claude Code: about two and a half hours, nearly 500,000 output tokens, about $65 → Codex: about one hour and 12 minutes, about 100,000 output tokens, about $20 Codex moved to 2-0. ## 3. Trail Latch Trail Latch was a portable kitchen for camping. Claude Code built the strongest product experience of the test up to that point. The page showed the kitchen unfolding, explained the prep deck, stove shelf, utensil rail, wash basin, dry bin, lantern arm, and side deck, then let me assemble a setup and see the price. The interactive latch sequence made the product easy to understand without reading a wall of copy. Codex took a more experimental approach. The opening composition was harder to read, and the scrolling journey gave my eyes too many places to go. Its secondary pages were much better. The setup builder, vehicle-fit page, products, and FAQs all felt cleaner than the homepage. Most visitors would judge the homepage first, so Claude Code won this round. → Claude Code: four agents, almost two hours, almost 400,000 output tokens, about $42 → Codex: one agent, about one hour, about 100,000 output tokens, just under $20 Claude Code had the better one-shot design, but Codex still left more room for efficient iteration. ## 4. Present and Clear Present and Clear was a public-speaking cohort built around clearer communication. Claude Code used a before-and-after idea and added interactive annotations that highlighted speaking problems. The concept was useful, but the page was too wordy and did not have enough visual depth. Codex crossed out "charisma" and emphasized a "clearer way through." The page also changed colors between sections and used stronger hierarchy with less copy. → Claude Code: four agents, about two hours, about 440,000 output tokens, about $42 → Codex: one agent, almost one hour, about 92,000 output tokens, just under $20 Codex won the design round. ## 5. North Ledger Studio North Ledger Studio was fractional CFO counsel for founder-led creative agencies. Claude Code's version felt like a long report. It had a lot of copy, very few visuals, and almost no change in pacing as I scrolled. Codex felt more like a real brand. It added navigation, a guided path through the page, specific agency metrics, and more color changes. It also created three clear ways to work together and built toward a focused call to action. The Claude page simply ended. → Claude Code: about two hours, about 500,000 output tokens, about $52 → Codex: about one hour, about 100,000 output tokens, about $18 ## The Open-Ended Test The sixth prompt removed nearly all creative constraints. I told each tool to invent the business, build the website, and show me what it could do with design. Claude Code created an acoustic-room concept around the line, "Your room has a note." It used ripples, material controls for the floor, walls, and ceiling, and a progress indicator along the side. The concept was interesting, but several elements were stretched and the page returned to the same problem I saw in earlier rounds: too much copy. Codex created "We Heard Tomorrow." It used a dark dynamic background, mouse-following effects, and large typography. The site turned a signal receiver into an interactive capture experience. The design was simpler, clearer, and more consistent. The resource gap was also the largest of the experiment. → Claude Code: almost three hours, roughly 330,000 output tokens, about $50 → Codex: about eight minutes and roughly $1.50 An open-ended prompt exposes the defaults inside the model and the harness. Claude Code tended to add more explanation and structure. Codex tended to cut copy and make a stronger first visual decision. ## Specific Prompts Changed the Result The most specific prompts erased the design gap. For build seven, I described nearly every piece of the website and told both tools exactly what I wanted. The outputs came back almost identical. The layout, copy, sections, and overall experience matched, with only small differences in things like icons. I repeated the test with an HTML report. Same result. You would not look at either pair and call one design significantly better than the other. Codex was still faster, cheaper, and more token-efficient. The design agent matters most when it has to fill in missing decisions. Once the prompt defines those decisions, capable tools tend to converge on the same result. ## The Final Numbers Across all eight builds: → Claude Code used 25 total agents. Codex used nine. → Claude Code took about 14 hours. Codex took about five. → Claude Code used almost three million output tokens. Codex used roughly 550,000. → Equivalent API billing was about $444 for Claude Code and about $100 for Codex. Codex won five of the six subjective design rounds. Claude Code won Trail Latch. The two highly specific builds were effectively ties. I am not treating that as a permanent verdict. I pay for both tools and use both every day. The models change, the harnesses change, and my opinion changes when I run another real comparison. The useful question is which tool fits the work in front of you and how much direction you are willing to provide. ## Wrap Open-ended prompts test a tool's judgment. Detailed prompts test its execution. This experiment made Codex look stronger at both design efficiency and resource efficiency, while Claude Code still produced one of my favorite individual pages. I walk through all eight builds in the full video. Link in the first reply.