Partly verified — 6 of 54 specificsA few specific details here couldn't be independently confirmed against the video. The overall summary is sound, but double-check exact numbers or names before you rely on them.
What this video is
⚡ a 11-minute video, readable in 60 seconds
Anthropic released Claude Sonnet 5.5 on September 28, 2026, the second model in the Claude 5.5 family. It runs 30%+ faster and costs up to 30% less for most work than Sonnet 5, while per-token pricing stays the same at $2 input / $10 output per million tokens. On Terminal-Bench 4.0, Sonnet 5.5 scores 70.6% versus Sonnet 5's 10.3%, and it is now the first Sonnet model to beat Pokemon Red using only screenshots. For anyone tracking Claude's model lineup, this is a capability jump at unchanged pricing, though the source notes Opus 5.5 still wins on large, complex projects and risky security requests still route to Sonnet 5.
New: Anthropic released Claude Sonnet 5.5 on September 28, 2026, the second model in the Claude 5.5 family. [00:02]
Key takeaways
New: Sonnet 5.5 runs 30%+ faster than Sonnet 5 and costs up to 30% less for most work. [00:02]
New: Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, versus Sonnet 5's 10.3%. [00:02]
New: Sonnet 5.5 is the first Sonnet model to beat Pokemon Red working only from screenshots. [00:03]
New: Sonnet 5.5 is the first Sonnet model to launch with cyber safeguards and fallbacks comparable to Opus 5's, since its cybersecurity capabilities match Opus 5's. [00:04]
+ 45 more takeaways
Added: Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks. [00:04]
New: On FrontierCode at High effort, Sonnet 5.5 scores 10 points higher than Sonnet 5 at the same setting, at about one fifteenth of the cost per task. [00:06]
New: On CursorBench, Sonnet 5.5's best score is within about two points of Opus 5.5. [00:06]
New: In head-to-head runs, Sonnet 5.5 batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs. [00:06]
New: Epic Games says Sonnet 5.5 cleared the quality bar expected of a higher-tier model, holding up on a system design audit and data flow review while managing tens of thousands of lines of code. [00:06]
New: On GDPval-AA, which tests 44 occupations across nine industries, Sonnet 5.5 scores nearly level with Opus 5.5 and about 400 points above Sonnet 5, and outperforms GPT-6 Solo on long-horizon knowledge work. [00:06]
New: Slack reports that without changing prompts, Sonnet 5.5 did better than Sonnet 5 on almost all offline Slack bot evals, in fewer steps and with about 14% fewer output tokens. [00:07]
New: In a live demo, Anthropic gave Sonnet 5 and Sonnet 5.5 the same request (a webpage with 400 birds flying in a swirling flock at sunset), starting at the same moment. [00:54]
New: Sonnet 5.5 wrote about two times more lines of code than Sonnet 5 in the demo. [01:10]
New: Sonnet 5.5 finished the demo task in about 10 seconds, before Sonnet 5 was done. [01:12]
New: The demo cost 4,100 output tokens for Sonnet 5.5 versus 4,600 for Sonnet 5. [01:27]
New: Finance company Balyasny tested the models on 2,441 real tasks: Sonnet 5 used about 497,000 tokens per answer versus about 121,000 for Sonnet 5.5, roughly 4 times less, with better answers. [02:31]
New: On Terminal-Bench, Sonnet 5.5 reportedly finishes 7 out of 10 real jobs on its own versus 1 for Sonnet 5. [02:39]
New: Opus 5.5, which costs twice as much, finished about 6.6 out of 10 on the same Terminal-Bench test, so Sonnet 5.5 edged it out. [03:01]
New: The new Sonnet does several steps at once, whereas the old Sonnet did things one at a time, letting it finish in fewer steps. [03:20]
New: App builder Base44 reported that Sonnet 5.5 rarely stopped mid-build to wait for an answer, so fewer builds get stuck. [03:07]
New: On Terminal-Bench 2.0 at Medium effort (the default in Claude apps), Sonnet 5.5 far exceeds Sonnet 5's best score for less than a tenth of the cost per task. [03:49]
Changed: Compared to Opus 5.5, Sonnet 5.5 does underperform but is significantly cheaper for less intense tasks. [03:54]
New: On a set of 44 real business jobs, experts judged Sonnet 5.5 basically tied with Opus 5.5, while Sonnet costs half as much to use. [04:10]
New: Streamer Chris Izatt raced four AIs through Pokemon Yellow simultaneously, with Sonnet 5.5 shown leading in the top-left position. [04:34]
New: Sonnet 5.5 built a game called Nahual, featuring a fire xolo that evolves into Xolcan, an axolotl-themed creature styled like Pokemon but in Spanish, from a single prompt. [04:57]
New: The Nahual game was 2,342 lines of code in a single file, including music, generated with no manual edits. [05:02]
New: On a benchmark of reading 100 charts, Sonnet 5 scored 16 out of 100 correct while Sonnet 5.5 scored about 62 out of 100. [05:17]
New: Sonnet 5.5 made a 28-second self-launch video with animated charts, glowing numbers, and sound design, reportedly built about two times faster than Opus 5.5. [05:29]
New: The video states Sonnet 5.5 writes more clearly than Sonnet 5, and testers said it feels more like working with a partner. [06:07]
New: A user (Gegam) had Sonnet 5.5 model a Roman legionary in Blender, calling the material work phenomenal, and said it beat GPT-6 Astra and GPT-6 Sol while landing roughly on par with Opus 5.5. [06:17]
New: Gegam says Sonnet 5.5 is very diligent, does not cut corners, and keeps polishing its work. [06:51]
New: These models are getting significantly better at checking and refining their own work, picking up render details like leather sheen, chainmail structure, and wood grain. [07:25]
New: Sonnet 5.5 shines at doing an entire coding job alone, hands off, and beats Opus 5.5 on that specific benchmark, while other benchmarks are relatively close. [08:56]
New: Sonnet 5.5 is literally half the cost of Opus 5.5 for very comparable results. [09:14]
New: Opus 5.5 still wins for big, messy projects that need a lot of careful thinking, according to Anthropic. [09:32]
New: Max effort on Sonnet 5.5 can be really slow, with one tester joking he would die of old age before his Max effort build finished. [09:34]
New: In Anthropic's locked-down safety testing, Sonnet 5.5 was the least likely of all its models to even poke at the walls. [09:56]
New: Creative coder Kevin Ngo suggests using Opus 5.5 for planning and Sonnet 5.5 for building to save costs. [10:01]
New: Opus can be used to check Sonnet's work, refine it, and send it back to Sonnet to execute the refinements. [10:08]
From $10 to about $7: a job that used to cost $10 on Sonnet 5 now costs about $7 on Sonnet 5.5, per Balyasny. [01:57]
From 10 seconds to about 7.5 seconds: answers that took 10 seconds on Sonnet 5 take about 7.5 seconds on Sonnet 5.5. [02:00]
From over $10 a job to under $1 a job: each job at Medium effort costs under $1 and still beats Sonnet 5's best score, which cost over $10 a job. [03:49]
Unchanged: Sonnet 5.5 is priced the same as Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads. [00:03]
Unchanged: Sonnet 5.5's biology safeguards are the same as Sonnet 5's, targeting a narrow set of high-risk requests while routine software development and most life sciences work are unaffected. [00:04]
Unchanged: Pricing per million tokens: Claude Fable 5.1 is $10 in/$50 out (5x), Opus 5.5 is $4/$20 (2x), Sonnet 5.5 is $2/$10 (baseline, same as Sonnet 5), and Haiku 4.5 is $1/$5 (half price). [09:06]
Unchanged: Risky security requests are still sent back to Sonnet 5 until Anthropic builds up guardrails for the new model, though normal coding and bug fixing are not affected. [09:37]
Added: Medium is the default effort setting, described as probably the best bang for your buck. [10:26]
Effective: Anthropic gave everybody one free usage reset that can be used anytime before October 22nd. [10:29]
Why it matters: at the same per-token price as Sonnet 5, Sonnet 5.5 completes far more coding jobs hands off and costs noticeably less per task, though Opus 5.5 remains the pick for large, complex projects and risky security requests still route to Sonnet 5.
New: Sonnet 5.5 made a 28-second self-launch video with animated charts, glowing numbers, and sound design, reportedly b ▶ 5:47New: Opus 5.5 still wins for big, messy projects that need a lot of careful thinking, according to Anthropic. [09:32] ▶ 9:33Unchanged: Risky security requests are still sent back to Sonnet 5 until Anthropic builds up guardrails for the new mode ▶ 9:39
Shown on screen — grab and go
PROMPTThe shared bird-flock benchmark prompt
Make a web page with 400 birds flying together in a swirling flock, like you see in the sky at sunset.
Shown on screen as the identical prompt given to both models; OCR rendered '400' as '4oo', corrected as an obvious digit/letter misread. shown at 1:27
QUOTESlack's customer quote on Sonnet 5.5
Without changing any of our prompts, Claude Sonnet 5.5 did better than Sonnet 5 on almost all of our offline Slackbot evals, in fewer steps and with about 14% fewer output tokens. When someone gives Slackbot a task, quality and speed are what matter most, and Sonnet 5.5 allows Slackbot to deliver better outcomes for users, faster.
Transcribed from the on-screen customer quote card, attributed to Slack; fully legible after restoring crushed spacing. shown at 0:06
QUOTEEpic Games' quote on coding quality
In Epic's early testing, Claude Sonnet 5.5 cleared the same quality bar you'd expect from a higher-tier model, holding up on a system design audit and a data flow review. The new model managed tens of thousands of lines of code for game [...] and delivered with less prescriptive prompting.
Transcribed from the on-screen quote attributed to Daniel Vogel, COO of Epic Games; one mid-sentence segment was too OCR-garbled to reconstruct and is marked with [...]. shown at 0:06
How this brief was shaped: What Changed · confidence Medium
Creator explicitly frames the video as 'here's what's actually new' comparing old Sonnet 5 to Sonnet 5.5, and OCR shows the official announcement copy with verbatim benchmark scores (Terminal-Bench 70.6% vs 10.3%, GDPval-AA) and pricing, while transcript runs a live before/after build race as proof of the change.
The lens sets this brief's structure, never its facts — every claim is held to the same citation and fact-check standard.