Creator Nate Herk pits the newly launched Claude Sonnet 5.5 against Claude Opus 5.5 across seven real work tasks, and the review lands on no single buy or skip verdict. Instead his rule is to use Sonnet 5.5 first when a task has a clear, objective definition of done, and reach for Opus 5.5 when you need a thought partner on a vaguer or more creative goal. Across the seven tests Sonnet won four rounds to Opus's three, with Opus typically costing roughly double per million tokens and only sometimes earning that premium through better motion design, copy, and taste. Herk stresses his round by round scoring is subjective and tells viewers to run the same test on their own workflows before choosing.
Verdict: The creator's rule is to use Sonnet 5.5 first when a task has an objective definition of done, and use Opus 5.5 when you need more creativity and a thought partner to help decide what done even looks like.
Key takeaways
Verdict: Across all seven test tasks, Sonnet 5.5 won four rounds and Opus 5.5 won three. [24:49]
Verdict: The creator calls his round by round judging subjective, for example awarding the PerkForm landing page round to Sonnet on that basis. [06:08]
Verdict: He tells viewers they must test each model on their own skills, prompts, and processes to find what works for them. [25:32]
For: Sonnet 5.5 is recommended for well-scoped everyday coding like fixing bugs and iterating on features, and for high-volume everyday development. [01:13]
+ 32 more takeaways
For: Sonnet 5.5 is recommended for polished documents, slides, and spreadsheets, and for well-defined repeated agent tasks like investigation, review, and drafting. [01:19]
For: Opus 5.5 is recommended for complex work requiring careful judgment, including long-horizon agentic coding and knowledge work, and for the hardest problems needing the most intelligence. [01:22]
Price: Sonnet 5.5 is $2 per million input tokens and $10 per million output tokens, versus $4 input and $20 output for Opus 5.5, making Opus roughly double Sonnet's cost.
Price: Overall, Sonnet 5.5 costs roughly half of what Opus 5.5 costs. [00:49]
Spec: Sonnet 5.5 scores 70.6% on the Terminal-Bench 4.0 agentic coding benchmark.
Spec: Anthropic says Sonnet 5.5 runs over 30% faster and costs up to 30% less than Sonnet 5 for most work.
Spec: Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
Spec: Across all seven tasks combined, Opus 5.5 totaled 3h 10m of active time and $66.67 in API cost, versus Sonnet 5.5 at 2h 34m and $40.56. [25:14]
Pro: Epic Games COO Daniel Vogel said Sonnet 5.5 cleared the quality bar of a higher-tier model in early testing, holding up on a system design audit and data-flow review while managing tens of thousands of lines of gameplay code and staying snappy on multi-hour tasks.
Pro: On the PerkForm landing page task, Sonnet finished in 28 minutes for $6.78 versus Opus's 52 minutes 49 seconds for $13.71, and the creator felt Opus's result was not twice as good despite costing twice as much.
Pro: For the YouTube resource guide task, Sonnet took 5 minutes 42 seconds for $1.47 versus Opus's 9 minutes 47 seconds for $2.87, making Sonnet about twice as fast and $1.40 cheaper. [24:18]
Pro: For the 90-day subscriber growth plan task, Sonnet cost $4.77 in 13 minutes 48 seconds versus Opus's $9.25 in 20 minutes 21 seconds. [13:54]
Con: On the PerkForm task, Sonnet's hero section used a generated fake can image instead of the real branded product image that Opus used.
Con: On the vague 90-day subscriber plan task, the creator felt Opus did much better than Sonnet despite costing more. [13:52]
Con: For the BrightPath financial model task, Opus was found cheaper and faster than Sonnet, at 28 minutes 57 seconds for $8.91 versus Sonnet's 30 minutes 29 seconds for $9.27. [17:25]
Pro: Opus 5.5 pulled in the actual real branded product images for the PerkForm landing page rather than generating a fake one.
Pro: The creator said Opus 5.5 is 'just so good as of now at motion design' and felt its output was more on brand in background and animation.
Pro: Opus is described as having better taste with motion and sound design. [09:45]
Pro: The creator felt Opus's Glaido video pulled an accurate quote from the real website matching a real quote from him, and that its copy was better than Sonnet's.
Pro: Opus generated an 18-slide investor deck for BrightPath Analytics including generated images and traction stats such as $6.42M ARR up 26% since January. [14:29]
Con: In one task, Opus ran about seven minutes longer and cost $6 more than Sonnet. [22:33]
Verdict: The creator awarded that round's win to Sonnet for being much cheaper, even though he felt Opus's single-prompt output was better. [23:05]
Vs: Nate states Opus wins the 90-day subscriber plan comparison. [12:55]
Vs: Opus ran for 20 minutes versus Sonnet's 14 minutes on that same task. [12:57]
Vs: For the Glaido motion video task, Sonnet took 53 minutes 34 seconds and cost $14.37, while Opus took 47 minutes 11 seconds and cost $20.55. [09:29]
Vs: Sonnet's plan launched How They AI on November 2nd rather than immediately in Q4. [10:41]
Vs: Opus's version ran as a Q4 90-day sprint from October 1st to December 29th. [11:45]
Reason: The creator states Opus is ultimately a better model overall, but says the point of the comparison is to show when and how to use each one. [14:01]
A Google Docs-style document showing the title and introduction of a guide: 'GPT 6 Astra $10,000 Stock Trading Challenge' with table of contents on the left sidebar.
It spent about 38 million input tokens more, as well as actually less output tokens, which is interesting, but it did cost about $16 more than Sonnet 5.
An investor pitch slide titled 'An $8M growth round' showing the raise amount, use of funds breakdown (in horizontal bars by category), and key financial assumptions for the closing.
Sonnet cost a little under $5, whereas Opus cost a little over $9.
For: Sonnet 5.5 is recommended for well-scoped everyday coding like fixing bugs and iterating on features, and for high- ▶ 1:11For: Sonnet 5.5 is recommended for polished documents, slides, and spreadsheets, and for well-defined repeated agent tas ▶ 1:19For: Opus 5.5 is recommended for complex work requiring careful judgment, including long-horizon agentic coding and know ▶ 1:23How this brief was shaped: Evaluation (review / comparison / unboxing) · confidence Medium
Transcript explicitly frames this as testing Sonnet 5.5 against Opus 5.5 across many task types with per-task verdicts like 'we give the win here to Sonnet', plus concrete price and speed specs (input/output token cost, percent faster/cheaper). OCR shows a built spreadsheet output (financial model workbook) confirming the demos are used as comparison evidence rather than a single tutorial.
The lens sets this brief's structure, never its facts — every claim is held to the same citation and fact-check standard.