Skip to playerSkip to main content
Business Enquires & Sponsorship: pritam.sahoo@gmail.com

Claude Fable 5.1 Is INSANE — But AI Companies Aren’t Telling You the Full Story

#claude #anthropic #fable

Category

🤖
Tech
Transcript
00:00cloud fable 5.1 is insane of course but ai companies are not telling you the full story
00:06so anthropic says cloud 5.1 fable is smarter and cheaper but here is the uncomfortable part
00:13independent testing suggests at maximum effort it can cost actually more per task than fable 5.
00:20so ai models are not really getting cheaper or are we being sold on wrong numbers that's
00:27the biggest question of the day we need some answers to it anthropic has released recently
00:32cloud fable 5.1 alongside mythos 5.1 if you are not aware mythos 5.1 is the more and
00:40more unrestricted
00:41version of fable 5.1 to select partners and companies and immediately internet did what
00:48internet always does gpt killer best coding model game over everything has changed and maybe whenever
00:56i see ai launch day hype happens like this now i have a different question forget the benchmarks
01:03what changes for you because the real story behind fable 5.1 isn't simply that cloud got smarter it's
01:10happening to economics autonomy and actual usefulness of ai and that is one part of this launch i think
01:17almost everybody is overlooking yes the model is extremely capable let's start with some benchmark
01:23numbers of course is not this is not a benchmark channel but i would like to highlight on anthropics
01:29terminal benchmark science evaluation fable 5.1 scored around 52.6 percent and fable 5 around 24.7
01:38that's not only a tiny increment improvement that's a significant jump honestly on terminal 4
01:45benchmark fable 5.1 score 55.8 percent compared to 42 percent for fable 5 and independent evaluator
01:54artificial analysis have also placed fable 5.1 extremely high on its intelligence ranking when
02:01running at maximum effort that's the key at maximum effort not medium high or something else so i don't
02:07want to manufacture controversy when there isn't any this appears to be a serious capable model
02:15and especially for complex reasoning coding agentic work and all those tasks very relevant in today's
02:22era but here is where i think ai industry has created a problem for itself which model is smartest i
02:28think
02:28that's increasingly wrong question honestly to me it's a stupid question i am going to say something that
02:34might upset a few ai benchmark fans benchmarks are becoming formula one statistics of ai they are useful
02:42they are interesting they tell us something very important but they don't necessarily tell you
02:47whether the model will make your business better that's the honest brutal answer a model can score
02:52two points higher on a benchmark and till be worse for your workflow because it takes longer it requires
02:58more human checking it doesn't integrate properly with your tools or it confidently gives an answer when
03:04it should have simply said i don't know and that's why i wouldn't choose claude gemini gpt or any model
03:11because somebody post a giant literacy board on x or other twitter.com your business doesn't get paid
03:19for benchmarks scores right it gets paid for simply outcomes and let me talk about the most important shift
03:28that is endurance what really caught my attention was with fable 5.1 is something different endurance
03:35when you are moving from ai where you ask can you write this mail towards where you say here is
03:42the
03:42objective go figure it out think about different how those two interactions are anthropic highlighted an
03:49example of from millennium where fable 5.1 reportedly identified the cause of extremely rare software crash that
03:56engineers have struggled for years mongodb says the model has worked on complex prototype across
04:02multiple days an engineer at ram described the machine learning workflow running unattended for about
04:0938 hours that's quite insane now these are some customer examples published by anthropic of course
04:16so treat them as early evidence not universal proof but the direction matters we are starting to measure ai
04:23not by intelligence but by how long it can remain useful without human constantly rescuing it that is
04:31where the things gets very interesting let me put one more angle to it cheaper ai can be a pricing
04:37illusion
04:38now let's talk about money because this is probably my favorite part of the story anthropic cut cash
04:44reading pricing from fable 5.1 dramatically around 75 percent and anthropic estimates typical workload could
04:52become 25 percent cheaper with highly agentic workloads potentially saving even more great but here is
05:00the catch independent artificial analysis testing around fable 5.1 at maximum effort cost around 3.76
05:08dollars per benchmark task but fable 5 was around 3.14 dollars per task so in that case the more
05:18efficient
05:19generation actually cost was about 24 to 20 percent more for task why because fable 5.1 generated some
05:28substantially more output tokens and this exposes something i think of a lot ai pricing headlines gets
05:35wrong pricing per token is not your ai bill it's like buying a car because petrol becomes cheaper without
05:42asking how much petrol the car actually consumes right it always about the task not about the tokens and
05:49if your ai agent thinks for 40 minutes call five tools generates thousands of tokens retries twice and
05:56validates its own answer so your cost is not simply 10 dollars per million tokens the real question is how
06:04much did it cost to finish the job right and one more crucial angle this is the metric enterprise should
06:09care
06:10about and i think this is where i think cios cto's founders need to change how they evaluate ai stop
06:17asking only how much does the model cost start asking what's the actual cost per successful outcome
06:24imagine a model a cost you one dollar it produces answer your employee spends 30 minutes checking it
06:30finds two hallucinations sends it back regenerates its checks it again and the ai bill may not be simply one
06:38more dollar again actually the business cost is much higher and model b cost let's say four dollars but
06:45it executes the workflow correctly sites its evidence use the appropriate tools and only needs five minutes
06:52of review model b looks four times more expensive on a pricing table but it actually be cheaper for the
06:58company the distinction is going to become more incredibly important as agentic ai spread across the
07:04enterprises and let me put one more angle to it ai may not replace your job because it could remove
07:09half of your for your workflow and here is the another part of ai discussion i think we have framed
07:15badly including me people keep asking will claude replace software engineers will ai replace consultants
07:22will ai agents replace analysts maybe that's how not disruption happens ai doesn't necessarily need to
07:29replace your entire job it is only here to remove six of ten steps inside your job think about a
07:36consultant research collect documents summarize competitors build a financial model draft slides
07:41check informations create recommendations prepare the meeting notes follow up today one person may
07:47spend a week doing all of that tomorrow an agent may complete six of the stages overnight while the
07:54consistent still exists and that to me is more realistic and potentially more disruptive than headline ai
08:01replace anyone and let me talk about the real war so when people ask me who is winning google open
08:07ai
08:08and tropic i think we are entering a phase that that question becomes much harder because the winner may
08:13not simply be whoever has the smartest model it may be whoever has the best combination of intelligence
08:19reliability tool use memory endurance security and finally the economics the best model in the world
08:25isn't useful if the companies can't trust it and the cheapest model isn't cheap if the employees spend
08:31hours correcting it that's the real ai war so here is my takeaway from fable 5.1 so anthropic appears
08:39to
08:39have made another major technical jump that's true but the more interesting story isn't claude bad gpt
08:45the headline will probably change again soon the biggest story is ai is moving from answer generation
08:51to work execution and once the ai being evaluated on completely tasks rather than clever answer the entire
08:58conversation changes benchmark matters less token pricing matters differently and human oversight becomes
09:05more important than not less because when ai can work independently for hours a small mistake can travel
09:11much farther before a human can see it so here is what i will do next forget the company benchmarks
09:18forget the launch presentations and give fable 5.1 and gpt 5.6 soul or the new model coming soon
09:27that is version 6 from gpt astra and gemini 3.8 is already out flash and test it with some
09:34real business
09:35assignment same prompt same document same objective and measure four things simple measure the four
09:41things simple quality hallucination time and more importantly cost per successful result and if you
09:47want to run that test comment real test below and send this video to someone who still thinks ai
09:53models by looking at benchmark screenshots because in the agentic era the smartest model may not be the
09:59model that wins and subscribe to the channel if you like this video hey if you've not met me my
10:04name is
10:04pritam i talk about ai latest tools latest hacks and i have been in this space for more than four
10:09years plus
10:11and see you in the next one
Comments

Recommended