What Fable 5 means for OMEGA users.
Anthropic took Claude Fable 5 to general availability on June 9. New flagship, a one-million-token context window, $10 per million input tokens and $50 per million output, live on the Anthropic API plus Bedrock, Vertex, and Foundry on day one. If you use OMEGA with your own Anthropic key, you were running it that afternoon. Not a beta channel, not a waitlist. The model showed up in your model list because the list comes from your key, and you assigned it to a slot and kept working.
That paragraph is the whole argument for model-agnostic architecture, so I want to unpack why it is harder to deliver than it sounds. Then I want to spend real time on the part launch threads skip: Fable 5 at frontier prices is a tool you aim, not a default you set and forget.
What same day actually means
OMEGA does not curate a model catalog. Bring-your-own-key means the set of models you can route to is read from the providers you hold keys for. When Anthropic flips a model to GA on their API, it exists for you at that moment. There is no OMEGA release between you and it, no integration sprint on our side that has to finish first. The same holds when OpenAI ships, when Google ships, when a new open-weight model lands in the MLX ecosystem.
It also means every surface at once. The slot you assign Fable 5 to is honored everywhere that slot is used:
- Main chat, including mid-conversation. You do not open a new thread to switch engines.
- Project Canvas, where a plan-as-nodes becomes real code and the planning nodes are exactly where a deep model earns its cost.
- Design Studio, where the copilot proposes reviewed, undoable edits on the canvas.
- Automations, so a scheduled run that fires at six on a Tuesday morning uses the model you picked, not the model we picked.
- Computer use and voice. Same selection, same key.
And if you route through a Claude subscription instead of a raw API key, that path works too. Subscription routing and BYOK land in the same slots, so the choice of how you pay Anthropic does not change what OMEGA can do with the model.
The integration queue you never agreed to
Here is the structural problem with most AI products: the vendor curates the model menu. When a frontier model ships, someone at that company has to evaluate it, tune their prompts against it, decide how the pricing passes through, and schedule a release. None of that is villainy. It is reasonable engineering caution. But it means there is a queue between you and every new model, and your position in that queue is not something you control.
The concession first: if you live inside Anthropic's own surfaces, June 9 was a great day. Claude had Fable 5 immediately, Claude Code had it in the terminal, and Cowork, their desktop agent preview that reads and edits files in folders you choose, runs on the same agentic core as Claude Code. Anthropic's first-party stack is the best possible day-one experience for an Anthropic model, and I am not going to pretend otherwise.
Local-runtime users know a version of this freedom already. If you run open-weight models through Ollama or LM Studio, nobody stands between you and a new release either. You pull the weights and go. I respect that path, and OMEGA speaks it natively through Omega-MLX on Apple Silicon. The limit is that frontier closed models never land there. Fable 5 is not something you can pull. For the closed frontier, a key is the only same-day door, and BYOK means you hold that key yourself.
But most people's work does not live inside one vendor's stack. The moment your assistant, your code tool, your design tool, and your automation runner are four different products, you are standing in four different queues, and they do not move at the same speed. Model-agnosticism is the decision to stand in zero of them.
Switching models should not cost you your memory
There is a quieter tax on trying a new model, and it bothers me more than the queue: your context lives inside products. ChatGPT's memory feature is genuinely good, and it is locked to ChatGPT. Claude Projects hold your documents and instructions, inside Claude. Each of these is useful right up until the day you want to try an engine the product does not offer, and then the price of curiosity is starting cold.
OMEGA puts memory a layer below the model. Four layers of it, plus the Brain, a knowledge graph that accumulates what you have told it and what it has learned across sessions. All of it is provider-portable, which is a dry phrase for something concrete: on June 9 you could be forty turns into a working session running on GPT-5.5, assign Fable 5 to that slot, and the forty-first turn runs on Fable 5 with the same conversation, the same Brain, the same project state. The new model reads everything the old one knew. Nothing gets re-explained.
That matters because throwing your real work at a model is how you find out what it is for. Thirteen days in, my own read is that Fable 5's gains show up on long-horizon work: plans with many dependent steps, and synthesis across more material than most models hold at once. You only learn that from your actual projects, and you only test on your actual projects if your context comes along for the ride.
The part nobody puts in the launch thread
Ten dollars per million input tokens and fifty per million output is real money, and I want to do the arithmetic in public, because surprise bills are how trust dies.
Every 100,000 tokens of context you send costs a dollar before the model writes a single word. Agent loops resend context on every turn, so a twenty-turn loop carrying that same context is twenty dollars of input alone. The one-million-token window is real, and I am glad it exists, and filling it once costs about ten dollars per call on the input side. Output at $50 per million means a long 5,000-token plan runs about twenty-five cents. None of these numbers is a complaint. Frontier compute costs what it costs. But a tool that hides this from you is not doing you a favor.
The wrong responses to that pricing are the two extremes. Never touching Fable 5 means the hardest work you do gets a weaker model than it deserves. Routing everything through it means paying deep-reasoning rates to answer "what time is my flight" from a note it already has. The right response is routing.
Aim the expensive model
OMEGA's standard engine is Neural-Fractal Agentic AI™ (NFA): work is decomposed into a tree of small, scoped cognitive units, and each unit is routed independently. I wrote up the architecture in an earlier post, but the short version is that decomposition is what makes frontier pricing manageable, because most units in a real task do not need a frontier model.
Routing runs through four slots, cheapest to heaviest, and you assign a model to each. A sensible June 2026 setup on a Mac looks like this: the cheapest slot runs a local model through Omega-MLX, which costs nothing per token and never leaves your machine, so trivial turns and quick recall are free. The middle slots take fast, inexpensive cloud models for tool-using errands. Fable 5 sits in the heavy slot and gets the turns that deserve it, like the plan for a multi-week build or the Project Canvas refactor that touches forty files. The same assignments govern OMEGA's 200+ agents and every Automation, and standing work is where discipline pays most. Dreams, the overnight self-improvement pass, is exactly the kind of always-on load you point at the local slot rather than the ten-dollar one.
Two properties of this system matter more than the slots themselves. First, assignment is entirely yours. There are no hardcoded fallback chains and no silent substitution. If you put Fable 5 in a slot, Fable 5 is what runs there, and if a provider has an outage OMEGA tells you plainly instead of quietly swapping in something you did not choose. Second, spend is governed. Keys are scoped per workspace, so a client project cannot draw down the key you use for personal work, and budgets put a hard ceiling on what gets spent. When you hand an always-on system an expensive engine, the cap is not optional.
No markup, no favorite
One more piece, because incentives explain products. OMEGA takes no markup on your tokens. Your Anthropic key is billed by Anthropic at Anthropic's prices, and the same holds for every provider we route to. We make nothing extra when you send a turn to Fable 5 and lose nothing when you send it to a free local model, which is exactly why the routing advice above has no angle: there is no margin pushing us to steer you toward the expensive engine. What OMEGA itself costs is on the pricing page, and it does not move with your token volume.
This is also why same-day support is a policy rather than a scramble. Fable 5 was the June case. Something else will ship soon enough, from Anthropic or from someone else, and the mechanics will be identical: it appears because your key can see it, and you assign it to the slot where it earns its price. Your memory comes along. The engines keep changing. The place your work lives should not.
Thirteen days in, here is my own scoreboard: Fable 5 is the strongest model I can hand a hard turn to, and the large majority of my turns still run on models that cost a fraction of it or nothing at all. Both halves of that sentence are the system working as designed.