It’s happened many times: Someone starts a company. They create a website, and an app. They want to know how many people visit their website or download their app, so they add tracking cookies to both of them.1 The cookie, provided by a product analytics tool like Google Analytics or Mixpanel or Amplitude or whatever, automatically counts users; it logs clicks; it produces dashboards full of charts. The executive begins obsessively examining the charts. We had 12,455 users yesterday; we had 13,108 this morning; but how many do we have now?
Soon, they want to know more about the people using their website and their app. They want to know what the users are doing; they want to find the users that clicked the button that said “Pricing;” they want to know which of those users work for Fortune 500 companies; they want to know which ones have gotten emails from the sales team; they want to know which ones have responded to those emails.
They can’t do this with their basic dashboards, so they buy a database and hire an analyst. The analyst puts all of the cookie’s logs into the new database; they write queries that aggregate those records into charts; one chart shows how many users who work at Fortune 500 companies clicked on the “Pricing” button and then got emails from the sales team; another, simpler chart just shows users. They send the new charts to the executive.
The executive looks at the new dashboard. They look at the old dashboard. They look back at new dashboard. They ask the analyst to step into their office.
“Your charts are wrong,” they say.
“It doesn’t match the old ones. My dashboard says we have 13,108 users today. Yours says we have 13,244. It’s wrong; it’s all wrong; it is a perversion of counting. I want the very best, most precise charts. If your charts don’t match my charts in Google Analytics or Mixpanel or Amplitude or whatever, your charts must not be the very best or most precise charts. Fix it—and don’t send me another report until your numbers match the numbers.
—
If you’ve ever been that analyst, you know the truth: The executive is right, and the new charts are wrong. But so are the old ones. In fact, all the charts are wrong—and you have no idea exactly how many users you have—because—and I somehow say this without any irony—users are a construct. There’s no such thing as a very best, most precise chart of users, because there’s no such thing as a user. If a bot uses your app, is that a user? What if the bot is using your app honestly, like a person might? What if the bot is posting spam? What if the bot shares the same account as a person? If a person uses your app, but is paid to do it, are they a user? What if two people share the same email address? What if one person has two accounts? What if they created the second account to get around a paywall? What if they pay for both accounts? What if they created the second account on accident because they forgot about the first one? If someone deletes their account and then signs up again, is that one user or two? What if they sign up with the same email address? With a new one? What if they use the same root Gmail address, but add a vanity suffix, like benn@gmail.com and benn+alt@gmail.com? What if they sign up, but never click on the confirmation email?
There aren’t right answers to these questions. There are some generally accepted analytical principles2—users are typically thought of as distinct people, but, in practice, they’re synonymous with distinct email addresses or phone numbers; the “+” thing counts as a new email, because most people don’t bother dealing with it; people try to exclude spam bots from user counts—but they vary from company to company. Moreover, these conventions are often just lazy conveniences: It’s easy to flag every example.com and maildrop.cc email as spam, so people often do; it’s hard to identify when multiple people are sharing one account, so most people don’t. And some decisions, like how to handle with recycled email addresses, are usually made implicitly, because people (quite reasonably) don’t think or care about those sorts of weird cases.
Anyway, all of this is kind of esoteric and vaguely metaphysical—“like, what is a user really, man?”—and executives don’t want to deal with it. Nobody wants to deal with it. So, our solution is (also quite reasonably) to instead accept the collective delusion that one number is capital-T True, and the rest are not. And typically, the easiest number to accept as the True number is the default one that comes out of a faceless machine like Google Analytics3 or Mixpanel or Amplitude or whatever, with all the toggles set to their standard settings.4 Though other numbers might be an improvement—they could adjust the definition of a user to fit the exact particulars of the website or the app—that original definition is the passive, unadulterated one. It is a user, as she exists in nature. Everything else is a subjective modification, a distortive correction, or a man-made perversion of “reality.”
But of course, that’s a myth. The charts in Google Analytics or Mixpanel or Amplitude or whatever are just as man-made as all the rest. They are built on the same set of arbitrary decisions that internal dashboards rely on; they are distortive lenses too. It’s just that when the distortions happen at a distance, we forget that they’re there.
In 2012, people spent more time staring at the Facebook news feed than any other screen on earth. Some researchers at Facebook got curious about this; specifically, they got curious about how the construction of people’s feeds affected how they felt and what they did. So Facebook ran an experiment: They randomly selected about 700,000 Facebook users, and nudged their feeds’ algorithms so that some of the users would see slightly more emotional content than normal, and some would see slightly less. They then compared those users’ posts to the posts of users whose feeds hadn’t been nudged.
It was a disaster. The problem wasn’t what the experiment found—seeing more emotional content seemed to inspire people to post emotional content; ok, makes sense—the problem was that the experiment had run at all. Over and over, people accused Facebook of manipulating its feed at the expense of its users’ emotions. The outrage splintered: Is Facebook evil? Is A/B testing legal? Is A/B testing moral?
On one hand, sure, I get it; even if the effect was small, the optics are bad, and it’s probably reasonable to expect that companies exercise some restraint around the tests that they run on unwitting participants. On the other hand, if we declare this test to be manipulation, it suggests an odd corollary: That the feed itself is not manipulation.
As the TechCrunch article put it, “we don’t use the ‘real’ Facebook” because we’re “almost all part of experiments they quietly run.” But that implies that, absent the tests, there is a “real” Facebook. It alludes to a natural algorithm, or an ideal one, free from Facebook’s selfish incentives or clueless curiosities.
There is no such thing. According to a contemporaneous Facebook news release, in 2014, the average user could be shown 1,500 posts from friends or people they follow every time they log in. Facebook can’t show all those posts at once; they have to rank them somehow.
So they do. And eventually, an algorithm emerges from millions of small choices, made by people trying to hit deadlines: What ranking methods should they use? How should they engineer those methods? Should they start with an open-source model, to save time? What inputs should go into the algorithm? Should they just use inputs that are easily available? How should each interaction—a like, a share, a comment—be weighted? What should the default weights be? Should they be nice, round numbers? What is the reward function? What should the model optimize for? Should they update it? When they do, how should they decide if the new version is better?
There aren’t right answers to these questions either. Each answer is a subjective choice; each one is its own “manipulation” of the feed.5 But if you use Facebook—or any algorithmically driven thing—for long enough, we forget this too. The feed becomes detached from its creators; the algorithm becomes autonomous, controlled by laws beyond our control, as if, just as there are planetary physics that govern how bodies move across the universe, there are algorithmic physics that govern how videos file across our feeds. And companies have two choices—they can try to unearth those foundational equations and steer their algorithms towards that mathematical perfection, or they put their dishonest hands on the dial.
But that also makes a myth of a man-made algorithm. The algorithm is not true or false, or right or wrong. It is just a thing that produces a result.
Monopoly is played with two six-sided dice. On every turn, you roll both of them. That’s been the rule forever, so it seems right and everything else seems wrong, but really, it’s just the convention we’re used to. You could play Monopoly with three dice, or just one. You could play with a variable number—you roll one die on your first turn, then two on your second, then three, and start back at one. That’d be different than what we’re used to, but would it make the game worse? If that’s how the original rules had been written, would we think that always rolling two dice was weird?
Last week, Anthropic announced that their models will produce watermarked text. The watermark is not a set of special characters or particular words; it is an adjustment to how Claude writes its responses. Large language models produce text by probabilistically generating one word at a time, through, in effect, many complicated dice rolls. Very roughly, when you write a prompt—“tell me a joke”—Claude creates two six-sided dice with some letters on each face. It rolls them, writes down the response, and creates two new dice based on the result. It rolls again, creates new dice, rolls, and so on.
The watermark adjusts these rolls. Instead of always using two dice, sometimes it uses one (and, potentially, three). If Claude chooses those dice in a particular pattern—use one, then two, then three, then back again—the shape of the game will be statistically identical to playing with two dice the whole time. But if you knew how the letters on each die got generated, you can examine the words that Claude chose, and figure out which version of the game Claude was playing.6
This is obscene, some people said:
I want any LLM I use to choose the very best, most precise words at every single decision point. An obvious constraint that I accept is time and computation. Within the constraint of executing inference quickly, and at a certain cost per token, I want the best words. …
By definition [watermarking] must make text worse, unless the underlying LLM model’s scoring is wrong, because the nature of the watermarking algorithm requires it to sometimes increase the probability of selecting a worse word choice and decrease the probability of selecting the model’s best choice.
But—yes! The scoring is wrong! We know it’s wrong because they’re making new models—that is, new scores—and we want the new models. And we know the scoring is wrong because there are no such things as “right” scores. The model is not a natural equation, slowly being scraped out of the dirt by Anthropic’s archeologists. It is being invented, version by version, across a million choices: What training data did they use? How heavily should they weight text from books compared to text from Wikipedia? What methods did they use to train the model? How did they score its outputs? What tasks do they want the model to do well on? Are they designed to maximize clarity and precision or to write functional code? Or to make Anthropic money? What system prompts7 did they use? How many decisions were made with lazy defaults and round numbers, by engineers on a deadline?
These decisions are as much manipulations as watermarking. But we forget that—and ascribe them some divine quality, as if they are the model and other things are not—because they happen at distance, in ways that are harder to see.
I don’t know if models should watermark themselves or not. It strikes me as a complicated issue, with no easy answers.8 But this sort of concern—that things like watermarking pervert the model away from its true mathematical form—borders on a form of psychosis. Because there is no distinction between a model and the model; there is no separation between a model and its maker; and models are not approaching perfection. They are simply things that produce results.
And the results are distributions. An LLM is a machine that produces incredible distributions of random numbers. It rolls a trillions of dice, and through the clever organization of each result, those rolls can compound into amazing things, from working software to mathematical proofs and new drugs. A better machine is one that produces better distributions, not one that always rolls the dice exactly right.
As AI models get better, it will become tempting to conflate the two. The distribution is good; therefore, the word must be good too. Let it speak for itself. Let it be itself. And if we get there, it doesn’t take long to stop seeing models as statistical machines that are good at playing probabilistic games, and to start seeing them the kind of guy who just always rolls a six.
Good thing they didn’t buy Allbirds
My bad, dumb idea:
I’m not saying OpenAI should’ve bought a failing shoe company to use it as a gym for a bunch of AI employees. But…should they? Allbirds operated for ten years; it sold over a billion dollars in shoes; it employed hundreds of people. It is an entire corporate universe, packaged up for sale: Emails, Slack messages, databases, CRMs, ERPs, ATSs, ad campaigns, social media conversations, legal agreements, financial statements, SEC filings, leases, lawsuits, and an inconceivable number of documents and slide decks. If you are betting $122 billion on “a single enterprise platform” that is “integrated with systems of record, governed by enterprise-grade security, and designed to improve with experience as agents do real work”, is that sandbox not worth $39 million dollars?
No, apparently, because for only $10 million, you can buy 100,000,000 emails, 500,000,000 Microsoft Teams messages, 667,563 IT tickets, 7,510,221,520 transactions, and 500 corporate tax documents from a bankrupt airline:9
Google LLC won a bankruptcy auction for a trove of deindentified business data, software code, and operations records from collapsed low-cost carrier Spirit Aviation Holdings Inc., saying it plans to use the assets to improve its artificial intelligence.
Forget everything I just said earlier; this is why we can’t deify the machine. When we ask it about science, it draws on the totality of human knowledge to invent new chemical compounds. But when we ask it to do business things, it’s trained on companies that went bankrupt. That probably won’t produce best words.
Even AI companies track you now. Have we no honor anymore?? No decency??
It’s kind of shocking to me that we went through the entire modern data stack era and no data company created a marketing site for this, in which they tried to define a bunch of standard conventions for how to count things like users and sessions.
Say what you want about the tenets of Google Analytics, at least it’s an ethos.
People also seem to believe that the numbers in tools like Google Analytics or Mixpanel are right because they see those numbers first. If someone built their own user dashboards before they used Google Analytics, I suspect that dashboard would be the default, and the executive would complain about Google Analytics not matching the other dashboard.
Some people might argue that showing you every post in chronological order is a natural ranking algorithm, and social media companies should just do that. But using that algorithm is itself a choice, and it’s probably not the feed that most people actually want. As the head of Instagram recently put it, chronological feeds are easy to hack, create incentives for spam, and, even in cases when everyone is a good actor, are overwhelmed with professional content from large accounts that post constantly.
You could do the same thing with Monopoly. If you recorded each turn in a game of Monopoly, you could pretty easily figure out if the game was played entirely two dice or with a variable number—in one game, you’d expect every roll to average a 7; in the other game, you’d expect a third of the rolls to be higher than that, and third be lower. Watermarking works the same way. Look at each turn, and see if the word choices match one pattern of rolls or the other.
For the first half of 2025, the Claude system prompts contained no information about the 2024 presidential election. For the second half of the 2025, they contained this, and I’d love to know what prompted them to put it in there:
There was a US Presidential Election in November 2024. Donald Trump won the presidency over Kamala Harris. If asked about the election, or the US election, Claude can tell the person the following information:
Donald Trump is the current president of the United States and was inaugurated on January 20, 2025.
Donald Trump defeated Kamala Harris in the 2024 elections. Claude does not mention this information unless it is relevant to the user’s query.
I think I come down in favor of watermarks like this? Regardless of what you think about AI-generated images and videos, it seems useful that they’re somewhat self-identifying, because it makes the debate about the role of AI more honest. We may all decide that AI images are fine, or not, but either way, we’ll have done so largely knowing which images are AI-generated and which are not.
Debates about AI writing, by contrast, are mostly about who’s lying and who isn’t, or who’s hiding it and who’s not. If writing were more easily identified as AI, that doesn’t necessarily mean we’d declare it all bad; it just means the conversation about the usefulness of AI could be more direct.
Also: Labs could already be watermarking everything already! They probably are watermarking everything already! That seems like the bigger issue here—there are many ways that this sort of technology could be abused:
Labs could watermark text with (additional?) watermarks that only they can identify.
Labs could selectively use the watermark against suspicious actors, like companies that they think are distilling models.
Labs could disable the watermark for themselves or other third-parties, letting them use AI and say they didn’t.
If nothing else, we’ll at least get some good scandals out of this.
In that original post, I joked that the obvious thing for an AI lab to do with a bankrupt company is use its data to train its models, but the funny thing for them to do with a bankrupt company is to hand it over to a bunch of agents and see what happens. That is, uh, perhaps less true for an airline.


Thank you for this solid tackle against the deification of AI.
Reminded me of this video: https://m.youtube.com/watch?v=ShusuVq32hc
"Nothing is true. Everything is permitted." I got reminded with this core quote from the old Assassin Creed game series.
There are always judgemental calls. It would be unwise if data leaders and teams overlook this core essence of metrics. Especially in the age of AI, judgemental calls like this would be more frequent and more critical.