Part of the harness engineering framework - on building agentic systems that outlast the model underneath them.
As of this morning, Fable 5 officially sits inside my Max subscription as standard, still at 50% weekly usage. For those who don’t know, Fable is the first of Anthropic’s Mythos-class models, the tier above Opus, built for (in their words) ambitious, long-running projects: work that plans across stages, delegates to sub-agents, and runs loops for literally days with minimal oversight.
In the world of AI, we tend to get weird months, and this took a genuinely strange month to get here, so I’ll give you the short version first. Then the question I actually want to put in front of you - because over the last month I ran the triage of how to translate multi-model usage, and here’s how I’d now recommend integrating AI seats (3rd party and local open-weight). I went through every scheduled task across multiple company agent fleets; the daily tasks, the marketing content transforms, CRM management, coding and bug fixes, feature roadmaps, the file audits, the pipeline sweeps - and defined which of these can be effectively scaled into a more automation format, and which still need frontier level thinking - Imagine this like a hiring question. How many times are you relying on the HoD for transactional tasks, and when does ‘scale’ relate to the execution of economically viable tasks?
The answer is almost none of the routine actions we are currently building towards in ‘efficiency AI’ are about frontier thought (apart from the framework to create the process transformation in the first place). Think of this like hiring a consultant. We are drafted in to define the problems, evaluate the opportunities, and architect the solutions. We are not running the operational tasks once its done.
How have we normalised frontier models?
Fable 5 launched on 9 June. Three days later the US government suspended it under an export-control order. It came back on 1 July with safeguards tightened, with only a one-week window before it was due to leave subscription plans (supposedly due to compute constraints) - this deadline that then moved from 7 July to the 12th, then the 19th, before Saturday’s announcement finally settled it. From today, it’s standard on Max and Team Premium at half of weekly limits. I think there are multiple reasons for this flip-flopping beyond the stated safety and compute concerns, most notably aligned to the launch and continued performance of other notable frontier model releases (Kimi 3, GPT 5.6, etc.) but that’s not the point.
Let’s focus on the model, its capability / purpose / and function in an augmented AI stack… A model judged sensitive enough for a government suspension is, just 4 weeks later, standard kit on consumer and business subscriptions - slightly tweaked, admittedly. That’s the speed at which the ‘extraordinary normalises’ now. It’s also five changes to architecture, consistency, access and pricing in just five weeks! If you’re in charge of AI transformation, how can you build a stack that aligns and interprets this level of change consistently?

The ‘mum AI’ question
Someone put it to me recently as the ‘mum AI’ question. It’s basically the 95% usage question - as a raw use function. Which model does your mum actually need? Her list: scale a recipe for six, reword a stroppy email, make sense of a council tax band, plan three days in Lisbon, or summarise X and Y. These are single tasks. They have been capable for years. Even the summarisation of complexity is more than capable with a single task model. Models from two years ago handle every item comfortably. Nothing on that list should get within a mile of a Mythos-class workload.
And the mum test isn’t really about mums. Swap her list for the queries flowing through most SMEs - summarise this meeting, first-draft this proposal, tidy this spreadsheet, pull data from the CRM, review and analyse the monthly sales figures, answer this customer, etc. Your operational query list looks more like her fridge list than you’d like, sorry. The overwhelming share of real-world AI use IS rudimentary. It’s useful, valuable, worth automating properly, but rudimentary - in the eyes of modern frontier model usage.
Businesses solved this decades ago, in staffing. You don’t put your most senior person on routine work; delegation, right? Resourcing to the task is one of the oldest disciplines in management. Flat-rate AI subscriptions have let everyone open up their capacity - but have muddied the waters with what they SHOULD be doing, because for a couple of years the frontier model cost the same as the small one and sat one dropdown away.

Where does Fable 5 earn its keep?
The routine lanes - briefing, transforms, audits - stay where they were, on capable mid-tier models with reasoning effort dialled to the task. They were never the frontier’s job, and moving them to Fable 5 would just change the invoice (token overspend), not the output.
Fable 5’s seat is in looking at the structure and strategy-grade work (it’s the consultant in the business): for long, multi-stage analysis where deeper reasoning changes the conclusion rather than the polish. It’s now the resident frontier tier in my stack for exactly that work, and nothing else. The test I used travels well for me and many SME businesses. If the extra reasoning genuinely changes the outcome, then pay for it. But remember, having the power is like having a £20k camera. If you don’t know how to use it, interpret the output, and work with the results, it’s just going to be ‘well reasoned noise’!
Two things make that workable in practice. The first is architectural: the harness outlasts the model underneath it, and my suggested profiles are all routing inside the harness that companies can actually own, rather than inside OpenAI or the hundreds of AI harness wrappers, so pointing a task at a different model is one line in a config, and moreover ‘one line to update’ when that model weighting changes for individual tasks. The second is that reasoning depth is now just a dial. Modern models let you set thinking effort and turn limits per task, which turns “spend reasoning” into an operational lever - the same discipline as putting expiry dates on AI skills, evals, and builds, applied to the spend and time itself. Every component in my architectures carry ‘review and eval’ triggers. So, model choice carries one too.
What does usage data show?
The largest study of real ChatGPT usage to date, published by OpenAI with the National Bureau of Economic Research, puts practical guidance, information-seeking and writing at nearly 80% of all conversations, with non-work use up from 53% to over 70% of messages and programming a small slice of the whole. Hold that against Anthropic’s own description of Fable 5 - “built for your most ambitious, long-running projects”, working “for days at a time” - at an API price of $10 per million input tokens and $50 per million out (double the Opus rate, ten times Haiku), and the overlap between what the ‘frontier tier’ is built for and what almost everyone asks it is pretty loose. The model isn’t at fault, user behaviour is though. The latest and greatest bias comes into play a lot here - but the way that Anthropic stuttered the release of Fable certainly influenced our novelty bias and consumption dependency of it too.

Match the tier to the task
So, do we need Fable 5? For the work it was built for, well yes - and when one of those jobs lands, I’m glad it’s now sitting in the plan permanently. But, for the everyday majority, we already had the model we needed, and for those looking to control token usage at an enterprise or innovation level, this harness and operational control layer is now MORE important than ever!
The integration question is simpler than the model cards make it look, and after the last five weeks it’s three layers of consideration rather than one. Who decides which model runs which work? Who decides which plan each seat actually needs, now that the frontier tier is standard on one subscription and metered on another? And who sets the effort level - because reasoning depth is a dial, and ‘token spend per task’ is a budget line - one judged by the difference in output validity, not necessarily the answer the agent gives. If the answer to any of those is “whatever the tool defaults to”, a procurement decision has been delegated to a dropdown.
Another eval and refinement stage is here. The model choice belongs inside the architecture loop now: made deliberately, reviewed every time a provider reprices or restructures token tiers. But 5 changes in 5 weeks says the reviews won’t be rare either!