This is why we can't have nice things
Do AI companies work yet?
I really thought it would be so simple:
So here’s an obvious prediction: AI will follow a nearly identical trajectory [as cloud computing]. In ten years, a new type of cloud—a generative one, a commercial Skynet, a public imagination—will undergird nearly every piece of technology we use. …
Just as cloud providers built out hundreds of utilities that are all underpinned by core services like EC2, I’d expect OpenAI to do the same thing on top of GPT and other foundational models. They already offer a chatbot, a speech-to-text service, and a text-to-image service. Surely, more utilities like these are coming: Text-to-video, video-to-text, text-to-audio, text-to-code, image-to-text, code-to-documentation, detection services to figure out if something was created or altered by an LLM, music generation, software generation, pipes between these services, and dozens more.
These products won’t be end-user applications, but developer tools. If you want to build on top of them, it’s a simple API call. Ask the ChatGPT API a question, and it’ll talk back to you. Send an image to it, and it’ll describe what it sees. Pass it a codebase and a desired change, and it’ll send you a new codebase with the requested feature. …
If this happens, public AI providers like OpenAI would become another backbone for the internet. Nearly every piece of technology will rely on their models. … But, as is the case for cloud providers, this critical infrastructure will be invisible to most people. … They’ll just come to expect that the products they buy to do the things that AI can do.
Back then, in 2023, it seemed like such a clean analogy. Twenty years ago, lots of companies wanted to store documents and host websites, but running server racks was complicated and expensive. So huge companies like Amazon, Microsoft, and Google built industrial villages full of server racks; now, nobody else has to buy their own server racks, and everyone just uses Amazon or Microsoft or Google to store their documents and host their websites.
Similarly, it seemed, everyone would want to use large language models, but training and hosting these models is complicated and expensive. So huge companies like OpenAI and Anthropic would build expert teams and technical infrastructure (and server racks) to do it; then, nobody else would have build their own models, and everyone would buy them from one of a handful of major model providers. And the market would support this arrangement, in which a small number of companies remain dominant, because:
…dominance is self-reinforcing. The bigger a provider becomes, the deeper its moat gets through an entrenched ecosystem, better models, and, likely, lower prices. And because training and running LLMs is very expensive (like building data centers is expensive), once a few AI providers separate themselves from the rest of the market, nobody else can catch up.
Ahaha, oops. From earlier this month:
Moonshot AI released a new A.I. model that appeared to narrow the lead held by well-funded American competitors.
Moonshot said that the model, Kimi K3, was the world’s largest open-source A.I. system, allowing anyone to use, modify and build on it freely. The company said that Kimi K3 performed as well as leading models from OpenAI and Anthropic at some key tasks.
…
The release of a free, open-source Chinese A.I. model that could rival the performance of costly, computing-intensive systems from Silicon Valley’s best-funded companies has reignited concerns over whether the industry’s enormous spending spree on data centers is justified.
We talked about this problem two years ago. Though developing a good AI model is really hard, it isn’t nearly as hard as building a cloud provider. Cloud providers are physical infrastructure that has to be bought, leased, manufactured, machined, shipped, soldered, assembled, electrified, secured, fortified, and maintained. There are limits on how quickly anyone can do those things:1 Nobody will ever drop a surprise cloud provider on AWS; no startup will ever “send shock waves through Silicon Valley” by coming out of stealth with 500 data centers.
An AI model is a file.2 A file can be posted on the internet; it can be developed quietly; it can appear overnight. Major AI providers’ biggest competitive advantage—their models, and the thin upgrade it offers relative to cheaper alternatives—is always one tweet away from disappearing.3 And this is especially true if distillation4 works, because it implies that every frontier model immediately becomes a tool for helping competitors catch up.5 The better economic analogy for a frontier model is a big-budget movie—it costs a ton to make, it makes a fortune for like a month,6 and is then relegated to some budget carousel on a “Previous releases” page. And research labs aren’t cloud providers, planting lasting infrastructure into the stubborn ground; they are expensive movie studios, in perpetual need of their next huge hit.
—
So if Kimi proves that some really popular movies can be cheap to make, what’s a lab to do? What is their business model?
An app for that
One answer is apps. (From that earlier post: “What, then, is an LLM vendor’s moat?...A better set of applications built on top of their core models?”) Maybe models have a short shelf-life, but if people prefer Claude Code over Codex, then they’ll become loyal customers of Anthropic’s models. Spend a zillion dollars with Anthropic, not because you like Claude, exactly, but because you like their coding app.
But that approach seems pretty fraught. Apps—or, you know, “wrappers”—are relatively cheap to build; there are already a number of popular open-source alternatives to tools like Claude Code. And it creates a teetering dependency, where the economic return from spending hundreds of billions of dollars on pioneering AI breakthroughs depends on the success of a hack day project.
A second answer is apps, but intertwined with the models. OpenAI is already optimizing some of its models specifically for Codex. It’s a soft cross-sell: Like Codex? It’s best with GPT-5. Like GPT-5? It’s best with Codex. But the more notable intertwining might be economic: Buy Claude Code, and get a discount on tokens at Anthropic. The importance of this intertwining became apparent when Anthropic tried to remove Fable, its best model, from its all-you-can-eat Claude Code plans,7 and require people to pay the standard per-token list price. The internet threw a fit; a bunch of people got wandering eyes for Codex; Anthropic said oops and changed their mind.
Moments like these, surely, raise questions about this whole situation. Business is booming—and maybe the market for this stuff is so big that nothing else matters—but these sorts of convulsions do make it all feel a little precarious. If botched pricing updates can quickly swing a market in another direction, how much of the market is driven by momentum, the vibes on Twitter, and things like “I asked Claude” starting to enter the lexicon?8 At some point, habit and reputation can be very deep moats,9 but we’re still pretty early in this whole AI thing, and early habits are easy to break.
So, again—what’s a lab to do, when the apps are just as easy to leave as the models?
A solution for that
I mean, here’s an easy answer: Make it hard to switch. Not emotionally hard—not because your product is just so good that people can’t imagine using something else—but, like, physically hard. Don’t sell simple APIs that can be replaced by a competitor’s in an hour; sell complicated ones, full of weird quirks and bespoke parameters and priced in impenetrable ways. Don’t sell apps that are easy to install and quick to set up; sell platforms that require layers of licensing and specialized deployment environments. Don’t integrate with other products in one click; integrate with them after three phone calls to a sales rep, two days of onsite implementation planning, and a mandatory certification program.
But you don’t say that. You say it’s custom pricing, designed to fit every customer’s financing needs. You say that it’s more secure this way. You say that you are committed to the success of your valued business partners, and that includes providing a platform that’s optimized for their complex enterprise needs. You say that you’re selling solutions.
You do this because this is how enterprise software works. Every product starts as a beautifully simple idea, and every company starts with the belief that if they just make something great, then everything else will take care of itself. It works for a while; everyone loves it; it is the noble hand of capitalism, making the world better, one seamless software experience at a time. And then, the company pivots to the enterprise, loses a deal to SAP or ServiceNow or Alteryx, and realizes that inscrutability is the point:
There is no killer feature; there is no sharp disruptive wedge. That’s why Alteryx makes almost a billion dollars a year, despite seeming like an easy target for startups. The maze of complexity is its defense. Trying to build a simplified Alteryx is how you lose to Alteryx.
AI has proven to be a pretty exciting thing for a lot of people, and that feel-good phase has lasted for a long time. But it won’t last forever; eventually, the pivot to the enterprise comes for the revolutionary stuff too. And then, maybe the labs will really become more like AWS or GCP—not just because they’re the backbones of the internet, but because using them will be need to be just as complicated.
Or, maybe not? As a somewhat extended endnote, I suppose there could be a more satisfying answer here.
My original post about OpenAI becoming AWS said that OpenAI would build a bunch of specialized models—“text-to-video, video-to-text, text-to-audio, text-to-code, image-to-text, code-to-documentation, detection services to figure out if something was created or altered by an LLM, music generation, software generation,” and so on. There are now music models, software generation models, detection services, and dozens more things like them; they just haven’t been built by the major labels, which have instead focused on building a handful of massive, general models.
So far, that decision had probably made sense; a really good general model can often be better at a specialized task than a specialized one. But you could imagine a point where that no longer holds. If a model has a good enough foundation—if it starts with the “intelligence” of Fable, GPT 5.6, or Kimi—training that model to be good at, say, drug development could yield much better results than trying to train the next generation of Fable to be good at everything.10
Or you could go a step further, and train the model to be good at whatever specific thing the customer wants.11 This is the strategy behind Thinking Machines, the AI lab founded by former OpenAI CTO Mira Murati. Microsoft CEO Satya Nadella also proposed something similar:
Every company is going to have to build what I think of as human capital and token capital. … This requires a new architectural approach where every business is able to build agentic systems that improve over time, while still retaining control over their IP. …
Companies need to turn their workflows, domain knowledge, and accumulated judgment into AI systems that improve with each use. Private evals should capture whether a model is actually improving against outcomes that matter to the business (not just external benchmarks!). Private reinforcement learning environments should let models grow stronger on real traces from inside the organization. …
This loop becomes the new IP of the firm. … And unlike most assets, it compounds. Every improved workflow generates better training signal, which accelerates the accumulation of tacit knowledge unique to the firm. The companies that build this early will have an advantage that is hard to replicate, regardless of any new individual model capability.
That’s another solution to the rapidly decaying value of what the labs are building—instead of just selling the model, sell the expertise that helped you build your model. Researchers keep pushing the frontier out; labs sell the frontier models directly, while making last season’s models available to customize, and offer their expertise to help people modify them. Or, as reader David Eubanks suggested in an email:
For AI agents to really replace human workers, the argument goes, they need to be more like humans: autonomous, with persistent memory, and accountable for their actions. If this came to pass, an AI company looks more like an HR placement outfit, matching agents to jobs. … This creates an evolutionary tree of AI models with different personalities and abilities.12
Right, if people are starting to give different “jobs” to different classes of AI models, why would they not eventually give those jobs to models that were specifically designed to do them?
For example, metal is heavy, and heavy stuff is hard to move.
Talking about this two years ago, I said this:
The billions that AWS spent on building data centers is a lasting defense. The billions that OpenAI spent on building prior versions of GPT is not, because better versions of it are already available for free on Github. Stylistically, Anthropic put itself deeply in the red to build ten incrementally better models; eight are now worthless, the ninth is open source, and the tenth is the thin technical edge that is keeping Anthropic alive. Whereas cloud providers can be disrupted, it would almost have to happen slowly. Every LLM vendor is eighteen months from dead.
Eighteen months might’ve been optimistic.
Distillation is the practice of using an existing model to train a new one, and makes it possible to train that new model much more cheaply than you could otherwise. It’s controversial, but whether or not it should happen is a different question than whether or not it will happen.
It’s the bike peloton or Mario Kart all over again: There are structural disadvantages to being in the lead, where your competitors can directly draft off of your efforts.
The third-highest grossing movie in history is Avatar: The Way of Water. The fifteenth highest is Avatar: Fire and Ash. Did you know these movies existed? Do you know which installment in the Avatar series each one is? Did you know that they are making at least two more of them? Can you name two characters in the Avatar series? Can you name two people who were in them? Is the most memorable thing about the entire Avatar series—which grossed almost $7 billion at the box office—the font?
Technically, these aren’t all-you-can-eat plans; they’re eat-a-certain-amount-for-a-flat-rate plans. But if you paid for the same amount à la carte—that is, if you bought the same amount of usage via the API as the flat-rate subscription plan allows for—the usage-based bill would be far higher.
Man, I am still curious about this though. What would happen if Claude said, “ah, our bad, Fable is now way cheaper!,” and then just started sending all their “Fable” requests to a server that was running Kimi? Would anyone even know?
“Nobody ever got fired for buying IBM.”
For example, during a recent interview on Lenny’s Podcast, Dianne Penn, who runs part of Anthropic’s research team, was asked why Claude writes the (awful) way that it does. She said the models have been trained to be good agents, but that’s come at the expense of good writing. And now, they’re focused on swinging back the other way, and making the models better at writing.
Which, sure, and it makes sense that coding agents would need to be at least passable writers so that they can explain what they did in remotely comprehensible ways. But it also suggests a different question: If models sort of waddle their way upwards, see-sawing between different skills as they get prioritized, could you fork the models? Could you just push really hard on the writing “skill,” and make something that’s genuinely good at that?
It’s not an AI model; it’s an AI solution.
If you’re interested in a far more expansive imagining of what this could look like, David’s written a series of stories exploring the idea.

Doesn't the studio analogy end more hopefully than it sounds? The studios that survived stopped living on single releases, they bought distribution and franchises, and the hits became marketing for the machine.