Skip to content
Artwork for the localhost
the localhost · Sunday · 44 min

the localhost:0003 | the ai hardware boom is beginning

Last week we argued the $25 token was dying. This week: that's exactly when the hardware boom starts. Frank takes the hot seat with a paradox. If intelligence keeps getting cheaper, we don't use less of it — we invent more things to do with it. More agents, more inference, more workloads. Which leaves every enterprise with one question they haven't had to ask before: which intelligence should we rent, and which should we own? Citi and Vercel are reportedly running open-weight models more than half the time. In June it was 29%. By late August, 53%. A new family of models out of the UAE ships six versions built for six classes of hardware — phone to server — with no quantization involved. NVIDIA's Hugging Face acquisition is official. And OpenAI shipped Astra. We also had four people instead of five, a guest from London, and for the first time in three episodes, a completely unanimous vote. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ CHAPTERS 00:00 If intelligence gets cheap, what do you buy? 00:47 Welcome back — who's here, who's at Disneyland 01:28 Laura Osborn joins from London 02:14 Neil, and the Bears/Seahawks problem 03:07 Chauncey, and the disclosure 03:43 The Rundown: NVIDIA and Hugging Face is official 04:25 Neil: what GitHub can tell us about this 05:48 Citi and Vercel pass 50% open weights 07:50 The three-tier model: endpoint, edge, cloud 08:56 Six models for six classes of hardware 10:22 Purpose-built beats shrinking things down? 11:21 What UK and European customers actually ask for 14:41 From cloud architect to local AI 15:18 OpenAI ships Astra 16:26 Chauncey: be careful what you point it at 17:51 Neil: code generation at the edge 19:33 Baseline: what should people actually care about? 20:20 Open weights are no longer philosophical 21:27 What's coming out of IFA 22:52 The medical school of the future 25:00 Why doesn't anyone know this is possible yet? 26:53 "We have more agents than human employees" 28:23 Tension of the Week: the hardware boom begins 30:27 Chauncey: fixed cost beats a variable one 32:39 Laura: I spent five years telling everyone to go cloud 34:18 Neil: when pennies turn into a million dollars 35:00 The brewery test 36:06 The data center next door 36:34 The vote 38:42 On My Device: the 36B gets the job 40:07 Chauncey: Copilot CLI, running locally 41:44 The Dream Team and the one-person virtual business 41:56 Laura: an on-device speaker coach 42:58 Neil: Mac vs the Beast, round two ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ WHAT CAME UP • NVIDIA acquiring Hugging Face is confirmed. NVIDIA says it stays open and that NVIDIA hardware won't be required. Neil's read: this looks like Microsoft buying GitHub — the same fears, and possibly the same outcome. His angle is that Hugging Face has been effectively off-limits to a lot of enterprises, and being inside a vendor with a security and compliance apparatus may be what finally gets those models sanctioned. • Citi and Vercel are reported to be running open-weight models for more than half their AI workloads — 29% in June, 53% by late August. Ten weeks. • The MBZUAI Institute of Foundational Models released six models spanning roughly 1B to 375B parameters, each trained for a specific class of hardware rather than quantized down from something larger. Frank ran the 36B on his dual-3090 workstation against Qwen3.8-27B: about 68 tokens/sec against 40, on a comparable qualifying profile. • Neil's three tiers: the endpoint, the cloud, and an emerging middle — departmental edge servers. Local doesn't grant you HIPAA or FERPA compliance by itself, but no round trip to the cloud is a smaller attack surface. • The economics nobody models: an insurance agent filing 50 claims a day, times 10,000 agents. Even at pennies per call, that's over a million dollars a year in workloads that could run at the edge. • Laura, five years a cloud architect, on doing a full 180: the constraint in Europe isn't just sovereignty rules, it's that cloud capacity is at its limits. • And Neil's field test of public sentiment, conducted on a bartender who asked him point blank whether he was anti-AI. The vote was 4–0. First unanimous verdict of the series. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ THIS WEEK'S CREW Frank Buchholz — host, independent Laura Osborn — guest, joining from London Neil Misak Chauncey Larsen Jacob Rhoades and Robert Henry are off this week. Jacob is celebrating his first anniversary. Robert was on a plane, sending Teams messages anyway. DISCLOSURE: Laura, Neil and Chauncey are employed by Microsoft. Frank is independent. All views expressed are their own, nothing discussed is unannounced or non-public, and none of this is financial advice. ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ localhost is a weekly show about running AI on hardware you own. Independent and unsponsored. New episodes weekly: https://thelocalhost.show There's no place like 127.0.0.1. #LocalAI #OpenWeights #AI #EdgeAI #AIHardware #Sovereignty #OnDeviceAI #AIPodcast

0:00-44:55

transcript

No transcript — this publisher did not publish one.

show notes

Last week we argued the $25 token was dying. This week: that's exactly when the hardware boom starts.

Frank takes the hot seat with a paradox. If intelligence keeps getting cheaper, we don't use less of it — we invent more things to do with it. More agents, more inference, more workloads. Which leaves every enterprise with one question they haven't had to ask before: which intelligence should we rent, and which should we own?

Citi and Vercel are reportedly running open-weight models more than half the time. In June it was 29%. By late August, 53%. A new family of models out of the UAE ships six versions built for six classes of hardware — phone to server — with no quantization involved. NVIDIA's Hugging Face acquisition is official. And OpenAI shipped Astra.

We also had four people instead of five, a guest from London, and for the first time in three episodes, a completely unanimous vote.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

CHAPTERS

00:00 If intelligence gets cheap, what do you buy?
00:47 Welcome back — who's here, who's at Disneyland
01:28 Laura Osborn joins from London
02:14 Neil, and the Bears/Seahawks problem
03:07 Chauncey, and the disclosure
03:43 The Rundown: NVIDIA and Hugging Face is official
04:25 Neil: what GitHub can tell us about this
05:48 Citi and Vercel pass 50% open weights
07:50 The three-tier model: endpoint, edge, cloud
08:56 Six models for six classes of hardware
10:22 Purpose-built beats shrinking things down?
11:21 What UK and European customers actually ask for
14:41 From cloud architect to local AI
15:18 OpenAI ships Astra
16:26 Chauncey: be careful what you point it at
17:51 Neil: code generation at the edge
19:33 Baseline: what should people actually care about?
20:20 Open weights are no longer philosophical
21:27 What's coming out of IFA
22:52 The medical school of the future
25:00 Why doesn't anyone know this is possible yet?
26:53 "We have more agents than human employees"
28:23 Tension of the Week: the hardware boom begins
30:27 Chauncey: fixed cost beats a variable one
32:39 Laura: I spent five years telling everyone to go cloud
34:18 Neil: when pennies turn into a million dollars
35:00 The brewery test
36:06 The data center next door
36:34 The vote
38:42 On My Device: the 36B gets the job
40:07 Chauncey: Copilot CLI, running locally
41:44 The Dream Team and the one-person virtual business
41:56 Laura: an on-device speaker coach
42:58 Neil: Mac vs the Beast, round two

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

WHAT CAME UP

• NVIDIA acquiring Hugging Face is confirmed. NVIDIA says it stays open and that NVIDIA hardware won't be required. Neil's read: this looks like Microsoft buying GitHub — the same fears, and possibly the same outcome. His angle is that Hugging Face has been effectively off-limits to a lot of enterprises, and being inside a vendor with a security and compliance apparatus may be what finally gets those models sanctioned.

• Citi and Vercel are reported to be running open-weight models for more than half their AI workloads — 29% in June, 53% by late August. Ten weeks.

• The MBZUAI Institute of Foundational Models released six models spanning roughly 1B to 375B parameters, each trained for a specific class of hardware rather than quantized down from something larger. Frank ran the 36B on his dual-3090 workstation against Qwen3.8-27B: about 68 tokens/sec against 40, on a comparable qualifying profile.

• Neil's three tiers: the endpoint, the cloud, and an emerging middle — departmental edge servers. Local doesn't grant you HIPAA or FERPA compliance by itself, but no round trip to the cloud is a smaller attack surface.

• The economics nobody models: an insurance agent filing 50 claims a day, times 10,000 agents. Even at pennies per call, that's over a million dollars a year in workloads that could run at the edge.

• Laura, five years a cloud architect, on doing a full 180: the constraint in Europe isn't just sovereignty rules, it's that cloud capacity is at its limits.

• And Neil's field test of public sentiment, conducted on a bartender who asked him point blank whether he was anti-AI.

The vote was 4–0. First unanimous verdict of the series.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

THIS WEEK'S CREW

Frank Buchholz — host, independent
Laura Osborn — guest, joining from London
Neil Misak
Chauncey Larsen

Jacob Rhoades and Robert Henry are off this week. Jacob is celebrating his first anniversary. Robert was on a plane, sending Teams messages anyway.

DISCLOSURE: Laura, Neil and Chauncey are employed by Microsoft. Frank is independent. All views expressed are their own, nothing discussed is unannounced or non-public, and none of this is financial advice.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

localhost is a weekly show about running AI on hardware you own. Independent and unsponsored.

New episodes weekly: https://thelocalhost.show

There's no place like 127.0.0.1.

#LocalAI #OpenWeights #AI #EdgeAI #AIHardware #Sovereignty #OnDeviceAI #AIPodcast