The Weekend a Model Got Banned (and Why I'm Glad Mine Lives Under My Desk)
Anthropic shipped Claude Fable 5 on June 9. Four days later it was gone. A weekend benchmark of two open models against the cloud tiers people actually reach for, and the case for getting serious about local.
Anthropic shipped Claude Fable 5 on June 9. Four days later it was gone.
Not deprecated. Not rate-limited. Banned. The US government handed Anthropic an export directive over a cybersecurity concern (a jailbreak that could help foreign users probe software vulnerabilities) and told them to cut off access for any foreign national. Anthropic couldn’t reliably verify nationality across API keys, consumer accounts, and enterprise integrations, so they did the only thing they could: they disabled Fable 5 and Mythos 5 for everyone, worldwide. A frontier model with a four-day lifespan.
If you’d built a workflow on it that week, the model didn’t get worse. Your access did. And it vanished by directive, on a timeline you had no say in.
That’s the whole argument for getting serious about local models, compressed into one news cycle. So this weekend I ran the experiment I’d been putting off.
The test: two open models, one gaming PC
I run a dual-GPU box: an RTX 5060 Ti (16GB) paired with an RTX 3060 (12GB), split across both cards in LM Studio. Nothing exotic. Parts you can buy this afternoon.
The test was a challenger match. Qwen3.6-27B has been my daily driver for a few weeks now: a dense model I trust for local work. Gemma 4 dropped recently, so I wanted to see whether the new release could unseat the incumbent. I gave both the same orchestration prompt: decompose a multi-step task, name a failure mode and a check for each sub-task, flag any ambiguity in the request, and return strict JSON with no markdown. The kind of job a planning agent actually does. Gemma 4 26B A4B is a mixture-of-experts model with roughly 4B active parameters, quantized to Q4. Gemma loaded at 64K context. Identical prompt to both.
Here’s how it went.
| Metric | Gemma 4 26B A4B | Qwen3.6-27B |
|---|---|---|
| Speed | 72.19 tok/s | 18.88 tok/s |
| Thinking time | 1m 11s | 2m 28s |
| Tokens generated | 5,574 | 3,231 |
| Valid JSON / schema | Pass | Pass |
| Markdown fences (told not to) | None | None |
| Parallelization reasoning | Correct (2 steps) | Correct (1 step) |
| Ambiguity caught | Surface (“urgent” undefined) | Structural (scope of a clause) |
| Failure-mode depth | Correct, thinner | Correct, one layer deeper |
The headline: Gemma was about four times faster and finished in half the wall-clock time. Both nailed the format. Neither leaked a markdown fence after being told not to, which is where a lot of models quietly fail. Qwen was the more careful thinker. It caught a structural ambiguity Gemma skated past, and its failure modes went one layer deeper. That tracks with the published benchmarks: Qwen wins on raw reasoning, Gemma wins on speed and efficiency, exactly what you’d expect from a sparse MoE versus a dense model.
Here’s what makes that result worth writing down. A few weeks ago Qwen was the best thing I could run. This weekend a brand-new model matched it on quality and quadrupled its speed, and I didn’t change a single piece of hardware. That’s the local-AI release cadence in one data point. New open models are landing every few weeks, each one resetting what a fixed box can do, and the curve is steep enough that “the best model my machine can run” is a different model month to month. The cloud improves on a roadmap you watch from the outside. Local improves on a roadmap you download.
So far, a clean local result. The interesting part is what happens when you put those numbers next to the cloud.
Now line it up against the cloud
The reflex is to assume a model running on a gaming PC is a toy next to a hosted endpoint. So I lined my two local numbers up against the cloud tiers people actually reach for: a fast model (Gemini Flash), a frontier reasoning model with its thinking turned up (Gemini 3.1 Pro), a fast hosted small model (Claude Haiku 4.5), and the model that got banned (Claude Fable 5, while it was still live). All figures are reported output speed.
| Model | Where it runs | Output speed |
|---|---|---|
| Gemini Flash | Cloud | ~230 tok/s |
| Gemini 3.1 Pro (thinking) | Cloud | ~130 tok/s |
| Claude Haiku 4.5 | Cloud | ~93 tok/s |
| Gemma 4 26B A4B | My desk | 72 tok/s |
| Claude Fable 5 (while live) | Cloud frontier | ~64 tok/s |
| Qwen3.6-27B | My desk | ~19 tok/s |
Read that table twice. Gemini Flash is the genuine speed demon, north of 230 tokens per second, and on raw speed the cloud wins without an argument. But look where Gemma lands. My mixture-of-experts model on two consumer GPUs out-ran a hosted Haiku endpoint, and it out-ran Claude Fable 5, the most capable model in this entire group when it was live. The frontier model was slower than my gaming PC.
That’s the part the pitch decks skip. You don’t rent the frontier for speed. Fable 5 and Gemini 3.1 Pro burn their time thinking: explicit chain-of-thought, dialed-up reasoning, and the latency that comes with it. What you’re paying for at the top is reasoning depth, and the benchmark that still separates the tiers is GPQA Diamond, where Fable 5 scored 92.6% and Gemini 3.1 Pro hit 94.3%. That’s the real cloud premium. Not tokens per second. Hard reasoning on hard problems.
Here’s the conjecture I’ll put my name on. For the large majority of what people actually use AI for (drafting, summarizing, extraction, classification, routine planning, reformatting, first-pass code) a frontier model like Gemini 3.1 Pro or Fable 5 is overkill. You’re paying frontier prices and accepting frontier risk to do work a model on your desk already does well enough. The 92% GPQA score is real, and it’s wasted on summarizing a meeting transcript. Most tasks don’t need a PhD-level reasoner. They need a competent one that’s fast, private, and always there.
So the honest split looks like this. For the planning, drafting, extraction, and classification that fill most of a day, my desk is already in the same field on speed and good enough on quality, with the data never leaving the building. For the genuinely hard reasoning, the small slice that actually stresses a model, the frontier is still worth the call. And the frontier is exactly the tier that just got switched off in four days.
It already fits in your pocket
The same weekend, I loaded Gemma 4 E2B onto my iPhone 16 through the Locally AI app. E2B is the tiny end of the family, built to run on a phone.
It worked better than it had any right to. I pointed the camera at plants, animals, and random household objects, and it identified them on-device, no signal required. I got voice mode running and asked it to explain Kubernetes. It gave a solid answer and mispronounced “Kubernetes” the entire time, which I found more charming than disqualifying. This is early. It’s also a multimodal model doing real vision and voice work, entirely on a phone I bought last year, with airplane mode on.
Once you’ve watched that happen, Apple’s strategy stops looking conservative and starts looking obvious.
Apple is betting the same way, with a hedge
At WWDC two weeks ago, Apple leaned hard into on-device intelligence. The new Foundation Models run locally, with Private Cloud Compute as the fallback for heavier work. They opened those cloud models to developers under two million downloads for free, pulling the infrastructure cost out of building AI features.
Two details matter for anyone reading the tea leaves.
First, the best on-device model is gated to the newest silicon: iPhone Air, iPhone 17 Pro and 17 Pro Max, iPads on M4 or later, Macs on M3 with at least 12GB of unified memory. Apple is selling the hardware that runs the model. The local-AI curve and the device-upgrade cycle are now the same business.
Second, the hedge. Apple confirmed its next-generation cloud Foundation Models were built in collaboration with Google’s Gemini. The company most committed to running AI on your device still rents a frontier brain for the hard jobs. Local where it can, cloud where it must. That’s not a contradiction. It’s the routing strategy everyone serious about this is converging on, including me.
Why the meter is the real story
There’s a reason the biggest companies in the world are suddenly counting tokens.
Meta employees burned through 73.7 trillion tokens in roughly thirty days on an internal tool they nicknamed “Claudeonomics.” Internal AI costs are running toward the billions, and Meta is now capping employee token budgets and building a system to monitor the spend. This is a company committing $135 billion to AI infrastructure, and even there the per-token meter became a problem worth governing.
Accenture saw the same pattern across its clients and turned it into a product: a service called TokenOps, built to optimize token consumption because the monthly Anthropic and OpenAI invoice went from a line item to a budget event. Their own forecast has global token usage multiplying twenty-four times by 2030, toward roughly 120 quadrillion tokens a month as agents take over the inference load.
Every one of those tokens is metered, priced by someone else, and subject to change. The price can move. The model can be deprecated. And as Fable 5 just demonstrated, the whole thing can be switched off by a directive that has nothing to do with you.
A model running under my desk has none of those properties. It won’t win a speed contest with Gemini Flash. It doesn’t have to. It has to be good enough for the job, private by default, and impossible to revoke. This weekend proved all three are already true on hardware I already owned.
The question I keep landing on: when the speed gap is this small and the control gap is this wide, what exactly are you renting the cloud for? For me the honest answer keeps shrinking. The hard jobs, and less of those every month.