Local AI is Waiting for its Raspberry Pi Moment.

Cheap silicon, open weights, and the ship’s computer in your pocket.

Local AI is nowhere near as mature as the noise suggests. We are at the ZX Spectrum stage of this thing. The enthusiasm is real, the community is real, the hardware is a compromise at every turn, and everyone involved is quietly pretending otherwise.

Before I get into what I actually want to see happen, it is worth being precise about what we mean by local AI, because the term is doing two very different jobs right now.

The first is Enterprise Local AI. A company decides it wants open weight models running on infrastructure it controls, either on premises or self-hosted in a cloud tenancy it owns. Cost matters, but it is a line item negotiated against compliance, data sovereignty and vendor risk. If the business case says spend six figures on GPU nodes, the business spends six figures on GPU nodes. This segment is fine. It will sort itself out because there is a procurement budget behind it.

The second is Personal Local AI. An individual wants capable models running on hardware they own, in their home, answering to nobody. Here cost is not a line item. Cost is the whole conversation. And this is the segment I care about, because it is the one that determines whether local AI becomes something normal people do or remains a hobby for those of us with more disposable income than sense.

The hardware market is actively hostile to you

You cannot talk about Personal Local AI without acknowledging that the broader AI boom has wrecked the consumer hardware market. Datacentre demand has hoovered up GPU supply and sent DRAM pricing through the roof. The knock-on effects land on everyone. Gamers are paying absurd prices for graphics cards. Anyone building a PC is paying more for memory than they were two years ago. The entire economics of personal computers have been distorted to feed the very industry that is now telling you to run models at home. The irony is not lost on me.

So what do you actually buy today if you want to run models locally?

Apple, reluctantly, is the answer

The honest answer is Apple hardware, and I say that as someone who is perfectly aware of what has happened to Apple’s prices. Unified memory is the reason. When your CPU and GPU share one large pool of fast RAM, a Mac Mini with 64GB can load models that a consumer NVIDIA card flatly refuses to touch, because discrete VRAM tops out long before your ambitions do.

The market has already voted on this. The OpenClaw explosion earlier this year turned the Mac Mini from Apple’s most forgettable product into the reference hardware for always-on personal agents, to the point where Tim Cook was telling analysts the Mini and the Studio would be supply constrained for months. Hermes Agent has driven a similar wave. Thousands of people have bought a small silver box purely to leave an agent running on it around the clock.

Is it the best platform on technical merit? No. macOS carries overheads that simply do not exist on a lean Linux install, and if you know what you are doing, a Linux box with the same memory will give you more of it back. AMD offers unified memory too through Strix Halo, and on paper it is a genuine alternative. But for the average person the comparison is not close. macOS is plug in, download, run. Linux is an evening of drivers, quantisation formats and inference server configuration, and most people have better things to do with their evenings. Apple wins by being the least annoying option, which is a very Apple way to win.

The appliance era has started

The newer development is the home AI appliance, and I think it matters more than its current sales figures suggest. Vora Aegis is the obvious example. A box you plug in, connect your accounts to, and it runs open models entirely in your home with essentially no setup. It is not cheap at around two thousand dollars, and the optional subscription tiers show the business model is still hedging its bets. But it fills a real gap. There is a large population of people who want the privacy and ownership story of local AI without ever wanting to learn what a GGUF file is. That population will grow, and this product category will grow with it.

Now the uncomfortable part

None of these options, not the Mac Mini, not a Strix Halo machine, not an appliance, will run the models you actually read about. If you want Kimi K3 at home, you are looking at 2.8 trillion parameters and heavy download north of one and a half terabytes. Even GLM 5.2 at 744 billion parameters, deliberately built to be the practical self-hosting option, is beyond anything you can reasonably put in a spare room. It is not impossible. People do build multi-node home clusters. But the cost makes any sane person wince, do the arithmetic, and conclude that the monthly subscription to a frontier lab is the rational choice. Which it is. That is the problem. The economics of Personal Local AI currently funnel you straight back into the arms of the cloud.

So we end up in a strange place where the open weight ecosystem has never been healthier, the models have never been better, and the average person has never been further from being able to run the good ones.

What needs to change

I do not think the answer is Apple shipping more memory or NVIDIA finding some generosity. The answer is a new industry. We need AI inference silicon designed from scratch for one job, running open weight models cheaply, manufactured in enormous volume on mature process nodes where the fabs have capacity to spare. Not repurposed gaming GPUs. Not datacentre parts with a consumer badge. Purpose-built inference chips, paired with commodity memory, at a price that makes a home inference box cost what a games console costs.

This is not fantasy economics. It is exactly what happened with every other computing wave. Mainframes gave way to minicomputers and gave way to the micro. The Raspberry Pi took what used to be a proper computer and made it a thirty-five pound impulse purchase, and an entire generation of tinkerers came out of it. Inference is a far more constrained problem than training, and constrained problems are where cheap specialised silicon thrives.

When that industry arrives, and I believe it will, it does more than serve local AI enthusiasts. It relieves the pressure that AI demand has put on GPUs and memory, and pricing across personal computers starts drifting back towards sanity. Gamers get their graphics cards back. Builders get affordable RAM back. And running a genuinely large model at home stops being a flex and starts being boring.

But cheap silicon in a box at home is only the halfway point, because the best Personal AI solution is not the one sitting at home waiting for you. It is the one you carry. A Mac Mini humming away in the study or an appliance in the cupboard under the stairs is tethered to a plug socket and a postcode, while your life happens at the school gate, on the train, in the meeting you should have declined. The moment you accept that, you arrive right back at your phone. And your phone means Apple.

Apple is sitting on the most interesting position in this entire market and, so far, has done remarkably little with it. They own the device you carry every waking hour, they own the silicon inside it, and that silicon already uses the same unified memory architecture that made the Mac Mini the darling of the local agent crowd. If Apple gets serious about running the growing crop of open weight models on the phone itself, not a distilled toy version, not a thin client to a datacentre, but genuinely capable models on the device, they do not just compete in the Personal AI market. They own it outright. Nobody else has the vertical integration to even attempt it.

And here is the part that makes it feasible rather than fantasy. The phone does not have to do it alone. Apple could ship a Mac Mini class AI appliance for the home, a proper inference box with serious unified memory, and pair it with the iPhone the way they already pair the Watch. The phone handles what it can on the device and hands the heavy lifting back to the box at home over a private, encrypted link that never touches anyone’s cloud. Your data stays on hardware you own, the big model runs where the memory lives, and the experience in your hand is seamless. No other company controls both ends of that link, the silicon in both devices and the operating systems tying them together.

Which is, when you think about it, exactly how Star Trek worked. The away team never carried the ship’s computer down to the planet. They carried a tricorder, and the tricorder talked to the ship. Steve Jobs spent his career taking science fiction hardware and making it mundane, and this was always the destination. A computer you simply talk to, that knows you, that answers to you alone, no server room, no subscription, no permission required from anyone. The tricorder in your pocket already exists. The ship’s computer is the unfinished part, and Apple is the only company positioned to build both ends and make them one thing.

Originally published on linkedin.com.