This video pits Fable 5.1 against Fable 5 by giving each model an identical prompt to autonomously build the same app, OpsFlow, while delegating engineering work to Opus and Sonnet worker agents. The presenter's verdict: Fable 5 won this particular challenge on cost, time, and efficiency versus the end result, even though a blind Codex evaluation scored Fable 5.1 higher overall at 9.1/10 versus Fable 5's 8.4/10. Fable 5.1's session cost $1,220.24 versus Fable 5's $519.64, and the presenter says he would not pay the roughly $700 premium for a similar end result, though he adds Fable 5.1 generally feels more efficient at understanding intent in everyday knowledge work.
Verdict: presenter opens by saying Fable 5.1 and Fable 5 built him the exact same app, and it wasn't even close [00:00]
Key takeaways
Verdict: based on cost, time, and efficiency versus the end result, presenter says Fable 5 won this particular challenge, even though Fable 5.1 scored higher overall in a blind evaluation [11:33]
Verdict: presenter says Fable 5.1 generally feels more efficient and better at understanding intent in day to day general knowledge work [11:41]
Verdict: presenter says he is not convinced Fable 5 was $600 worse and proposes testing both models on an equal budget [09:43]
Spec: the app both models built is called OpsFlow, a node based workflow builder [00:00]
+ 17 more takeaways
Price: Fable 5.1's session cost $1,220.24 total, using Opus 57%, Sonnet 40%, cache hit 98% [09:10]
Price: Fable 5's session cost $519.64 total, using Sonnet 80%, Opus 12%, cache hit 97% [09:10]
Spec: Fable 5.1's session ran on localhost port 5321 for 36 hours and used about 40% (404,000 tokens) of its context window [11:07]
Spec: Fable 5's session ran on port 4382 and used 260,000 tokens, about 26% of its context window [11:24]
Spec: a blind Codex evaluation scored Fable 5.1 9.1/10 weighted versus Fable 5's 8.4/10 [08:00]
Spec: category scores for Fable 5 versus Fable 5.1 were visual design 8.1 vs 9.1, first run ease of use 7.6 vs 9.2, workflow authoring 9.1 vs 8.7, dry run experience 8.1 vs 9.3, validation and safety 9.4 vs 9.3, and accessibility 8.3 vs 8.8 [08:12]
Pro (Fable 5): deeper authoring controls including more condition operators such as regex, existence checks, list, and negation, plus live preview, an autosave timestamp, import and export controls, and negative test presets like Missing severity [08:37]
Pro (Fable 5.1): stronger information hierarchy, a more readable canvas, visible payload presets and preview, working manual approval that correctly fails the run on rejection instead of falsely claiming mitigation, clear visual states, downloadable run logs, and a better accessibility foundation with skip links [08:48]
Con: one of the two versions' trigger node only allowed connecting to one downstream node at a time, producing a connection refused error reading API error spike detected, already has an ongoing connection [04:20]
Con: Fable 5 and Fable 5.1 use different JSON schemas, so a workflow exported from one fails to import into the other and the canvas stays empty
Con: presenter prefers Fable 5's zoomed out text rendering because it stays a consistent size, while the other version's text grows and looks cheap when zoomed in
Vs: Fable 5.1 took about a day and a half to build versus Fable 5's half a day [00:10]
Vs: both models got identical prompts to own strategy, planning, delegation, sequencing, quality standards, and final acceptance rather than personally perform engineering, research, debugging, testing, or visual design [00:47]
Vs: both were told to primarily use Opus workers for architecture, product, design direction, difficult problem solving, and independent reviews [01:01]
Vs: both were told to primarily use Sonnet workers for implementation, research, testing, debugging, and iteration [01:07]
Reason: presenter says Fable 5.1 cost about $700 more than Fable 5 for a similar end result, and he would not pay that premium [12:06]
Not only did these apps look and feel very different, but one cost me $1,200 and one cost me $500.
How this brief was shaped: Evaluation (review / comparison / unboxing) · confidence Medium
Opening explicitly frames this as a head-to-head test of Fable 5.1 vs Fable 5 building the identical app from the identical prompt, with cost ($1,200 vs $500) and time-to-complete (a day and a half vs half a day) stated as the comparison metrics, which is a classic verdict/specs evaluation shape. The 25% slice then shows the creator reviewing a workflow run's output as supporting evidence, and OCR confirms a real automation app (OpsFlow) with trigger/condition/action/approval nodes actually being executed and inspected.
The lens sets this brief's structure, never its facts — every claim is held to the same citation and fact-check standard.