The Hardware Stood Still
Two years, one frozen memory ceiling, and a model that got nearly five times smarter anyway
For two years, the best laptop you could buy barely improved at running AI. The model it could run got nearly five times smarter regardless.
That is the claim from Clement Delangue, who runs Hugging Face, the site that hosts most of the freely downloadable AI models. He did the arithmetic. Between May 2024 and May 2026, the most expensive MacBook Pro on sale stayed at 128 gigabytes of memory. That number matters because memory is the hard limit on how large a model your machine can hold. It did not budge.
The model did. The smartest freely downloadable model you could actually fit on that laptop went from a score of 10 to 47 on one widely cited industry benchmark. Delangue’s framing: “Local open-weight AI on a laptop has been improving more than twice as fast as Moore’s Law!” Moore’s Law being the decades-old rule of thumb that computer chips roughly double in power every couple of years.
This article was reviewed and hand-edited, using curated sources I hand-picked, drafted by AI using personalised style guidance. I find it useful to keep in touch with the dozens of X.com bookmarks I come across every few days. Maybe you do too.
What that 4.7x actually buys
A benchmark score is not the same as a laptop being four times more useful. Scores compress a lot of messy reality into one number. But the direction is real, and we have said as much in recent issues: the local stack keeps closing the gap with hosted AI services.
What is new is the cause. The gains did not come from hardware. Apple has not given laptop buyers more memory since early 2024. Every bit of that improvement came from better models and better software squeezing more out of a fixed ceiling. In Delangue’s own words: “the hardware ceiling barely moved.”
What people are actually running
A chief executive can do the headline math. The thread that tells you the ground truth came from someone else.
Hermes, who posts as @HermesAgentTips, asked the plain question: what are local AI people actually running right now? He listed the obvious candidates, from a DGX Spark, the rig everyone supposedly has under the desk, down to years-old gaming graphics cards repurposed to run AI. His own words: “everyone talks like they have a DGX Spark under the desk, but I’m curious what the real setups look like”.
The answers were all over the place.
One reader runs a DGX Spark, a Mac Studio desktop, and is adding a graphics card.
Another has two of those Sparks plus a Mac Studio, waiting on software to pool them into one.
One is on a single mid-range gaming card.
One is still on an old Intel MacBook paired with a budget graphics card from years back.
One built a desktop around a current high-end gaming card.
There was no consensus and no standard rig. The “DGX Spark under the desk” is partly a status symbol: when one reader mentioned owning two, Hermes replied “NOT ONE BUT 2 SPARKS very nice”.
You might not need the big rig to do useful work
The most useful reply came from a reader posting as leetllm. They run a single mid-range gaming graphics card with 16 gigabytes of onboard memory. On it sits a shrunk-down version of Qwen, a freely downloadable model from Alibaba. Their verdict: “4-bit qwen 3.5 fits perfectly in 16gb vram and is shockingly capable. you really don’t need a massive rig.”
That cuts against our own recent coverage. We have repeatedly flagged that serious local AI still demands an expensive Apple Silicon machine or a multi-card workstation. leetllm’s setup is neither. It is a single card a lot of gamers already own.
The catch: “shockingly capable” is one person’s read on their own workload. Capable at what? For how long a task? leetllm does not say. A 16-gigabyte card runs a small model well and hits a wall the moment you ask for something bigger. But the point stands. The entry price for usable local AI is lower than the thread’s framing suggests.
Why bother at all? A reader running a professional-grade card put it simply. They are moving their core models to local versions because “I never want to see another rate limit or API error ever again :)”. That is the honest motivation for most of these setups. Not speed, not secrecy. Escape from the throttling and outages of paid online services.
Where the money still buys something
If the ceiling is fixed, the frontier is pooling machines.
Salvatore Sanfilippo, who built the Redis database and posts as antirez, was just gifted a top-spec MacBook Pro with this year’s Apple chip. What he wants to do with it is the interesting part. He plans to split a single model across two laptops, this year’s machine and last year’s, so they run as one. Another reader in the Hermes thread is doing the same with three machines, waiting on the software to stitch them together.
This is the actual high end of local AI right now. Not one enormous box. Several ordinary ones, lashed together, because the single-machine ceiling is moving too slowly for this impatient crowd.
🔮 Prediction
The 128-gigabyte ceiling holds through at least the end of 2026. Apple has shown no urgency to lift it, and it has not needed to. So the next year of “local AI got better” stories will be almost entirely about software and models, not new hardware. The new capability will come from pooling cheap machines, the antirez approach, rather than buying one expensive one. What would change my mind: Apple or Nvidia shipping a real jump in top-end memory at a consumer price. Nothing in these bookmarks suggests that is close.
What are you actually running? I would like to know, same as Hermes did.
The ceiling is still 128. Everything worth watching is happening underneath it.

