A River of Tokens

I don’t think anyone is totally certain which layer of the AI stack the value accrues to. But people are trying to capture as much of it as possible anyway. Everybody.

Think of tokens flowing downstream on a river. It’s like the oil industry. You get the oil somewhere, you refine it somewhere else, you sell it to an airline, and the airplane flies people. Some of the value of a gallon of oil is left behind at each step as it passes through the economy. Tokens work the same way, starting from the data before the model is even trained. Someone is harvesting data. Someone is training a model. The tokens pass through the GPUs and the data centers, then go to an application, and then a customer grabs them and makes some decisions. Follow the river and ask the same question at every step: who is going to keep the margins?

Start upstream, with data. Last year the labs spent a billion dollars each acquiring it. A while back, data meant tokens for pretraining: a bunch of text. Labs were buying old books and destroying them to get the text, distilling the fundamental knowledge of humanity and the internet. But training over the last year has moved towards reinforcement learning, and data for reinforcement learning is a different type of data. It’s an environment, a playground on which an agent is given tasks to accomplish. Imagine a task hard enough that the agent gets it right 1% of the time. You wait until it gets it right, grab the result, and reinforce that behavior. That’s how you make a model like Claude better at finance: you have to teach it finance. The capability a general purpose model has is proportional to where the lab chooses to train it.

That billion dollars flows to data vendors. Scale is the well known one, but there’s a long list of small vendors making $50 million a year. They’re printing money. And they rent the work out in turn: hourly workers, 1099s, usually outside the US, grinding through tasks in accounting, finance, chemistry.

Keep going downstream and you hit the chips. People think NVIDIA is going to accrue a lot of the upside, and NVIDIA still has a very large dominance. But now you have a bunch of domain specific chips, and OpenAI building its own chip, and you start wondering: is NVIDIA really going to capture a lot of the margins here? Then you look at data centers. Google, Microsoft, and Amazon are building them; they know how to build them, and they have the financing. But all these neoclouds are building too. So maybe it’s not that either.

Then you reach the models, and the models are very good. But you do see new entrants. You have open source. You have Grok suddenly back in the running, SpaceX a real contender, and a bunch of companies in China releasing large open source models. And the capabilities are interchangeable: through OpenRouter you can send the same request to any model. If the model becomes a commodity, you can still have the best model, but maybe people don’t use it exclusively, and competition can catch up at any time. That will totally compress the margins.

Anthropic’s revenue is real, and it’s going to continue growing. It’s also disorganized. I’m working on a research grant, and the other day I looked and I had spent $9 on something. That is not how anyone will operate in the steady state. This is a messy, disorganized period, and it’s less clear what the revenue looks like once people have more financial discipline, or if there’s a retraction in the market.

So who is going to keep the margins? Maybe the applications. Think about it this way. In the legal space, you have Legora, and who are the other guys? Harvey. Basically two. Legal is pretty big, there are two real players, and if one of them wins, it’s probably going to have a lot of stickiness in that market for a long time. At the general purpose model level, you have Anthropic, OpenAI, Google, now SpaceX, plus the open source models from China. Way more competitive at the model level, less competitive at the vertical level. Of course the model market is larger, because a model can be applied to everything.

But whose revenue is durable? Anthropic selling an API in the back end? Or the vertical company that owns the customer relationships? Any time that company feels it’s paying too much for one API, it can switch to another, or build its own model. Cursor fine tuned an open source model, and then SpaceX acquired them. If more companies build their own models on top of open source models from China, and those companies own the relationship with the customer, that’s a threat to Anthropic and OpenAI. Whoever directly interacts with the customer is what matters. Who is the front end to this intelligence? If you’re the front end, and you have a way to eat the margins down the stack, by having your own model, or by being large enough for your own data center, then that’s bad for everybody below you.

So why would Anthropic or OpenAI ever get into finance or legal? There are real reasons they wouldn’t. Anthropic is very oriented towards science and hardcore technology. They outsource data to vendors because they consider dealing with data beneath them. They don’t want salespeople. They don’t want customer service. They just want to be doing research. That’s not who they are. It’s not their DNA. And if they’re convinced they will have the best model forever, and everybody will use it, and they will keep the margins, then the vertical applications are effectively distribution for them, and they don’t care. Every vertical AI company becomes a channel: it charges a margin on top of your tokens, and you sit back as the token tax man, keeping a lot of that revenue. Why would you go into a vertical and deal with a bunch of salespeople?

Except they already own one vertical: developers. Claude Code and Codex. Eighteen months ago, Cursor was the distribution channel for coders and the labs sold the models through it. Then it was determined that, no, actually, we should do the CLI tools and sell directly to developers. That really dampened Cursor’s growth. It was a phenomenal exit still. And you can make the same case for any of these verticals. Say a customer pays $10 per million tokens. Some of that goes to the application layer, some to the model, some to AWS, some to NVIDIA. If the labs realize the share going to the application layer is high enough, they will want part of it. And right now they have the leverage: they are the most highly valued companies in the stack, and everybody depends on the models.

But coding was the easy case. Claude Code just took off on its own. You can hope lawyers and people in finance pick up Claude Cowork the same way, and hope that’s sufficient. It seems less likely to me. It’s one thing to organically displace a competitor the way Claude Code did. It’s a completely different thing to build the legal tools in house and just beat that market, because there are a lot of layers to a vertical company even when the technology overlaps: the packaging, the relationships, the customer success. Going to conferences and having booths. Forward deployed engineers inside all these companies. Wining and dining people. Hiring a bunch of former executives and being part of that network. Running a whole industry. And Anthropic doesn’t want to be a company selling to legal firms. It wants to get all the big verticals. Which means buying not one but 20. What does that look like? Are you a conglomerate of a bunch of companies? Do you take controlling interests, like a private equity firm? Do you just invest in a bunch of them? Who knows. Either way, the goal is the same: control the supply chain of tokens, and guarantee that all that demand goes to you.

Some of it is already happening. OpenAI bought a small health records company. Google has a health operation; even Apple has a health operation. Google was in talks to acquire Mechanize, a data vendor. Meta, Scale. Google bought the airline data out of a bankruptcy. NVIDIA is opening data centers. Everybody is going after everyone else’s stretch of the river.

Will they do it? Does it make sense for them? Is it a distraction, given how fast they’re growing? Is it premature? Perhaps. These companies are about to IPO, and there are too many verticals to go and take just one. But look at the biggest sectors: health, finance, legal, the military. Any job where people get paid a lot of money to read and write is going to be a hot market. How the industry consolidates is an open question. How many of these industries can one company run? I don’t know. AI allows you to scale more. So maybe they can.

At the very end of the river, the token dissipates, like the exhaust of your car. Dumb. That token did all the work it was going to do. Everybody upstream is trying to capture as much of that value as possible first. This is how the game is being played.