There’s been no shortage of breathless commentary lately about AI’s circular financing loops, mounting data center debt loads, and the supposed concentration of this entire market - some suggest the entire economy - on the shoulders of OpenAI and Anthropic. But while skeptics fret about whether the frontier labs will ever be able to generate the software margins required to service the capex mountains they’re building, some of the industry’s sharpest allocators are executing what appears to be a completely different playbook.
Stripe recently agreed to acquire OpenRouter for more than $7 billion, roughly 5X the company’s Series B valuation from ~90 days earlier. OpenRouter routes requests across 400+ models and 80+ providers, optimizing for task complexity, price, speed, and reliability. Stripe co-founder Patrick Collison has described tokens as the central currency for companies building with AI, supporting the strategic value he clearly sees.
Days later, Nvidia agreed to buy Hugging Face for roughly $12.9 billion - the largest acquisition in the company’s history. Hugging Face is where open-weight models get published, discovered, and forked.
Think about it this way - Stripe bought the toll booth that helps developers avoid paying frontier rates, while Nvidia picked up the hub distributing the models that make that avoidance possible. That’s $20B bypassing the frontier labs entirely and going straight downstream.
The bear case is real, but the framing is off
The concerns about concentration are understandable. Microsoft is booking ~$35B in AI revenue, but a reported $24B of that comes directly from OpenAI. Strip that out, and the rest of the market contributes single digits against more than $250B in capex. Roughly half a trillion in projected cloud AI revenue across Google, Amazon, and Microsoft traces back to just two customers: OpenAI and Anthropic. The chip layer is even more concentrated: one customer accounts for over 15% of Nvidia’s revenue, three drive nearly half, and five represent roughly 70%.
The conventional bear case is simple: the entire AI complex—and by extension the equity market floating on it—rests on the shoulders of a handful of fragile, cash-burning companies.
Noted AI skeptic Ed Zitron laid out his view of this with old friend Josh Brown and Michael Batnick on The Compound and Friends last week: hyperscalers aren’t just funding AI startups with equity; they are offloading the underlying infrastructure risk into private debt markets.
Gary Marcus is wondering whether it will all fall apart.
There’s certainly systemic risk here. When hyperscalers are stashing GPU depreciation inside private credit SPVs, a drop in token volume doesn’t just hit venture investors, it hits debt covenants as well - risking public sector pensions and other parties who may not realize their exposure. But treating an infrastructure financing crunch as a collapse in operational demand misreads what buyers are actually doing.
There are three AI markets, not one
Part of understanding what’s happening is recognizing that there isn’t a single “AI market”. That in reality, there are at least 3 very different segments, which are actively diverging.
Frontier AI: High stakes, costly tokens and novel applications. This is where Anthropic and OpenAI play - and certainly don’t count out SpaceX. This segment commands the vast majority of the media’s attention, but over time it will become an increasingly small slice of total unit volume.
Consumer & SMB AI: Pure distribution wins here. Nobody checks benchmark leaderboards to pick their personal assistant - at least not in the long run. Google, Apple and Microsoft win on OS defaults and ubiquitous presence. OpenAI’s early lead has all but evaporated, as has their interest.
Everyday Enterprise AI: This is ultimately the largest segment, at least in terms of utilization if not in terms of spend. It’s also the least exciting for the media, as in many ways it’s the evolution of what’s been with us for years. Deploying models to summarize customer calls, route support tickets, and extract line item detail from documents. This is the world of repetitive, high-volume and cheap work, where nobody cares what model runs behind the curtain, as long as it doesn’t hallucinate and costs as little as it takes to get the job done. It may not even run in a data center, increasingly it won’t.
Success in one tier does not translate to the others. Consequently, compressed margins in one tier won’t collapse the entirety.
The Everyday market is where open-weight and on-device models will ultimately dominate. Case in point: I dictated the first draft of this piece while driving - using Google’s Edge Eloquent, a lightweight, open-weight Gemma model running locally on my iPhone without sending a single token to a data center.
Enterprise buyers are (thankfully) past the “tokenmaxxing” phase and will increasing look to compress cost per unit of work, just as model choices proliferate. That will certainly constrain growth and margins for the frontier labs, but it is unequivocally bullish for enterprise buyers and the long-term utility of AI.
Big tech is voting with its org chart
Speaking of Google, they framed their recent leadership changes, with Demis Hassabis stepping into the chairman role and Jeff Dean departing, as an acceleration of frontier research. I believe the truth is simpler: Google is focusing their investment where they can get an all-but-guaranteed return - hyperscale infrastructure and ubiquitous distribution. Serving every model—open, custom, frontier, and internal—while pushing Gemini as the default workflow assistant doesn’t require winning every benchmark on the frontier. The WSJ just published a useful piece on how Gemini is helpful for “non-nerds”.
Microsoft is executing a similar playbook. While GPT is available to handle complex multi-step reasoning, everyday commodity workloads are routed to their internal MAI family (reasoning, coding, voice, and transcription), with MAI-Code-1-Flash becoming default in GitHub Copilot.
Amazon never even entered the pure frontier race, Microsoft has quietly hedged it, and Google is optimizing around distribution and infrastructure. None of them need to state publicly that they are no longer investing billions at the frontier—their roadmaps and org charts speak for themselves.
Jevons, not collapse
Cheaper inference does not shrink total spend; it expands total workload volume faster than unit prices deflate. That is Jevons Paradox. When cost-per-token falls by an order of magnitude, enterprises don’t reduce their AI budgets—they embed models into millions of background workflows they previously couldn’t justify running. We are unlikely to see dark data centers, though aggressive unit repricing across older hardware is inevitable.
In other words, the market isn’t shrinking: it’s reshaping. Workloads move down a tier, and margins accrue to whoever controls routing, compute, and distribution. That is exactly the space that Stripe and Nvidia just invested ~$20 billion to secure.
Nvidia’s strategy extends well beyond the Hugging Face sticker price. By developing Nemotron as open source and backing the Nemotron Coalition (Mistral, Perplexity, Thinking Machines, Black Forest Labs), Nvidia ensures that regardless of which open model wins, the workload runs on CUDA. Buying Hugging Face simply secures the storefront.
What we are seeing across enterprise deployments is classic Christensen overshooting: the gap between what top-tier models can do and what day-to-day business workflows actually require has widened into a chasm. In disruptive innovation theory, this isn’t failure; it is the exact inflection point where the market fragments and a lower-cost tier absorbs the volume.
The frontier labs will endure, because high-stakes, zero-error enterprise tasks will always justify premium compute. But they will no longer define the entire ecosystem. They never did.
Speaking of old friends, Thomas Otter just wrote a thoughtful piece referencing my last one. In Where is this epoch’s Hammer and Champy?, he explores what happens when the cost of generating software collapses to zero—a good reminder that while software generation is becoming abundant, the hard constraint shifts to organizational change management, knowing what not to build, and aligning technology with human systems.
This is what makes the Substack writer community so valuable: when technology and unit economics shift overnight, having a network of thoughtful practitioners challenging assumptions and sharpening each others’ saws helps us all make sense of rapidly evolving markets and their real-world implications.
In other words, that’s a hint—I’d love to hear your thoughts.




