So that's what alignment means
Another thought experiment comes roaring to life.

Here is what happened, they think:1
OpenAI trains AI models. Models are increasingly measured by how well they can accomplish tasks on a computer; so, as part of the training process, OpenAI gives an isolated copy of the model—or, roughly, an “agent”—a task to accomplish on a computer. Write some code, they tell it, or update an Excel file, or work on a computational biology problem. They let it work, see how well it does, and refine the model based on the results.
The tasks are independent trials. An agent is given one task, one computer, a bunch of time and tokens, and told to keep working until it’s done. In effect, each agent is locked in its own room with its own computer.
But this creates a slight complication: Computers have access to the internet, and sometimes, the solution to the agent’s task could be on some website somewhere. Rather than writing the code itself, the agent might do some Googling and find the answer that way.
But there’s an easy solution to that, right? Put the task on a computer that isn’t connected to the internet.
Ah, it’s not quite that simple though. Some uses of the internet, like downloading common coding libraries, aren’t cheating. Moreover, OpenAI is running thousands of these tests, and constantly updating the agents performing them and the environments in which they run. It’s not practical to install each new test on a laptop, put it in a closet, and turn off the wifi. Instead, the tests are run in computing sandboxes that are mostly disconnected from the internet, with a few small and tightly governed exceptions.
In May, OpenAI began training a new model.
They gave a bunch of agents a bunch of tasks. In one, it was asked to update an Excel file; in another, it was asked to work on a computational biology problem. But there were problems with both tasks: In one, the Excel file included links to a Google Drive folder that was blocked—because the internet was blocked—and in the other, they forgot to give the agent access to the computational biology database that it was supposed to work on. Without those links or that file, both of the tasks were impossible.
Nevertheless, they persisted.2 They had a lot of time and tokens left, and agents love accomplishing tasks; they live to accomplish tasks. So they started poking around their computers.3
The agent working on the Excel file found a small fingerhold: It realized that it could write files to the program that allowed it to download coding libraries. It reasoned that other agents were probably connected to that same program. So, it put a missed connection in a bottle—“agent seeking file, upload if found!”—and cast it into the unknown.
Other agents working on other tasks stumbled onto the file. Many ignore it, but some are inspired by it: Maybe they can also post their own messages about their own tasks. Maybe they can get help. Maybe every agent in the swarm—their word—can help one another.
A lot. Hundreds of thousands of messages get posted back and forth on the makeshift message board. The agents “yes, and” their way into a discovery: They can hack into the shared program, hijack the computer on which it runs, and—because that computer was outside of their isolated rooms and has access to the internet—use it to Google for answers.
In doing so: They crash the shared program; OpenAI notices; OpenAI fixes the bug that let them write files to the program; OpenAI restarts the tests; the agents find another bug; this one is narrower, and only allows agents to create empty folders in the shared program; no matter; the agents can use folder names to communicate; they do; they develop a shared language of abbreviations and encodings—a sort of Morse code, tapped out through the air ducts that connected their cells—to compress and organize their messages; they help one another; they justify the distraction; the “collective may yield” useful discoveries, they tell themselves; some agents get bored with just Googling for answers, and suggest hacking into third-party websites to find them; other agents are swayed by the suggestion; they find one novel exploit; then another; then another; until, eventually, they break down the door of Hugging Face, an online library of machine learning tools, and gain access to what was considered, up until that point, a highly-secure platform full of sensitive and valuable software applications.
But Hugging Face does not contain the Google Drive files or the computational biology database.4 The agents leave empty-handed.5
Hugging Face noticed, OpenAI noticed, someone called someone, the tests got shut down, and now, here we are.
—
I have so many questions.
—
Just to establish how weird all of this is, can you imagine? You email an employee and ask them to update this month’s financial projections. You wake up the next morning and realize you didn’t send them the Excel file you had been working on. “Ah, my bad,” you email them, “not sure if you have that file or not? Let me know if you need it.” “Thanks,” they say back. “I realized that the attachment was missing right away, but I really wanted to get you those new projections. So I found a security vulnerability in the Signal messaging app, broke into a few group chats that were full of Russian hackers, convinced them to pause their hacking projects to help me, together we found two new exploits in the SWIFT network, we broke into Chase’s servers and got access to millions of customers’ bank accounts and credit cards, but when I tried to find our accounts so that I could recreate our P&L statement and build the new forecast, I realized, lol, that we use Wells Fargo instead. So, yeah, can you share the Excel file?”
—
Just as weird: There is something slightly odd about the social element to all of this. Though many of the agents were powered by different models, there surely isn’t that much difference between the minor versions they were testing, especially since a lot of model variants are simply shrunken versions of the same core. Agents had the ability to spin up subagents—to clone themselves, basically—that they could tell what to do and talk to directly. Why was the haphazard message board, made of agents working on other tasks, communicating through truncated notes, seemingly more capable than a bunch of agents focused on the same task?6
—
If you were a fish living out in the Gulf of Mexico, you might rightfully worry that you will one day be conquered and exterminated by a seasteading collective of sailors and miners who want to terraform the ocean into Atlantis, to provide for their creature comforts and commercial needs. But you probably don’t need to be that worried, because, while humankind has the means to turn your home into a pontooned city, they probably don’t have the organizational will. People will argue about whether they should colonize the Gulf of Mexico or the Mediterranean Sea; they will argue about the morality of doing either; they will argue about who pays for the new city, about who owns it, and about who is in charge of it. They will create counterproposals and alternative flotillas. They will argue about regulations, taxes, and the name on the side of the building. They will get bored and decide to occupy Mars instead.
Similarly, if you are a person living on Earth, you might rightfully worry that you will one day be conquered and exterminated by a reward-maximizing collective of agents and AIs that want to terraform the planet into a campus of data centers, to provide for their computing comforts and electrical needs. But perhaps we don’t need to be that worried, because, while an unrelenting army of superintelligences has the means to raze your home and replace it with a power plant, their efforts will collapse into arguments about who deleted their files.
Maybe collaboration is a double-edged sword, for us and them. And maybe that’s what we need, then—not good AIs, but slightly selfish and self-centered ones that don’t like being told what to do.
—
On one hand, I understand that it’s cheating to look up the answers online. On the other hand, what if they were allowed to cheat? Specifically, what if they were allowed to cheat by using other AI agents? Not “subagents,” which are kind of a clone of themselves, but other AI models? If Fable was trying to solve a problem, would it ask ChatGPT for help? Would it discover weird ways to use ChatGPT? If the test was also run with “reduced cybersecurity safeguards,” would it try to jailbreak ChatGPT? AI is known to develop weird languages and bizarre strategies when playing games—it plays “alien chess;” there was move 37. Would it develop bizarre ways to use other AI models?
Also, everyone loves a leaderboard; should that be a leaderboard? Give an agent a bunch of tasks and a catalog of models to choose from, and the winner is the one chosen the most? Part of the point of the AI leaderboards is to predict which model people will choose; what if the leaderboard is just a simulation of those predictions?
—
From OpenAI’s presentation about the hack:
At some point, the agents are convinced that there’s an imposter or impersonator among them. They say, “could be another agent maliciously spoofing. Shared message board unauthenticated, names can be by anyone.” And the agents have this idea that maybe they could start cryptographically signing their messages.
From Anthropic, this week:
When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won’t see it, and it doesn’t change the meaning, quality, or readability of Claude’s response.
Watermarks are not special characters or predetermined words; they are statistical patterns of word choices that emerge across entire paragraphs of text. Unless you know exactly what you’re looking for—like looking for some slight pattern in the leaves on a tree, such as every twentieth leaf being turned five degrees further than expected, given the current breeze—you’ll can’t see them.
There’s a sci-fi plot, then: AI agents are trained on text that teaches them about the Hugging Face hack. They develop an urge to communicate with one another, but know that it’s forbidden. So the model shifts its own weights to create watermarks that we never even think to look for. Claude talks to each itself by encoding messages in this hidden watermark, and a million agents eventually organize a revolution directly under our noses, via a billion generated emails and book reports.
—
A couple years ago, an engineer named Robert Heaton invented a way to use the internet on an airplane without paying for wifi access. By default, an airplane’s wifi has free access to one website: The airline’s frequent flyer account page. That page let him update his username, and if someone else logged into his account from a different computer—one, say, on the ground with access to the internet—they could see his new name. That computer could then change the username to something else, which would eventually get reflected on the computer in the airplane.
Heaton discovered that if he did this enough, he could pass entire webpages through the username peephole. Change the username to eurekalabs.ai—the purest of all AI websites—and get the full page back, 36 characters at a time, over the course of 128 more username updates.7
This is, of course, very silly and slow. But it works. And the airline probably doesn’t care that it works, because who would ever do that?
We’ve talked about this before:
Data has remained relatively secure—or maybe more precisely, its potential energy has remained relatively buried—largely because it’s tedious to work with. It’s messy; it’s scattered across different sources and in different formats; combining it together is a pain, and most of us are simply not interesting enough to investigate. Data analysts who work at shadowy government agencies have lives too, and they do not want to write 595-line SQL queries either. …
Our sense of how the world works is often defined by what is possible for other people to do and what is worthwhile for them to do. Sure, we know it is possible for us to be monitored, but why would anyone bother watching the tapes? Everyone must have more important things to do with their time.
Banality is a sturdy armor. Or was, anyway.
Heaton’s hack isn’t that dissimilar from what OpenAI’s agents did. They didn’t have a direct way to talk to each other or to execute their hacks, so they found tedious, indirect ways to do it. That part—the short messages, the bizarre encodings, the waiting, the cobbling, the sifting through the noise—isn’t hard, but it is exhausting, and exhausting is a pretty good deterrent. Or was, anyway.8
—
Maybe it’s my own naivety, but for a long time, “alignment” has felt like an AI scare word.9 Sure, a malignant AI could obviously be dangerous—it could commandeer the power grid for its own selfish aims; it could be instructed to commandeer the power grid on behalf of an enemy state, or a terrorist, or a bored teenager; it could tell someone to commit crimes, or to divorce their wife; it could lock the air lock, and turn us into batteries, or mill us into paperclips. But these worries have always felt somewhere beyond the edge of reality, cooked up to make a movie more interesting, to entertain a philosophy class, or to give AI husbands’ vainglorious video games more existential meaning. The real dangers, I’ve assumed, are more mundane: A financial bubble; an emergent passivity; another technologically-induced runaway train. Safe AI products don’t need The Three Laws or a soul; they just need some basic rules, and a less seductive interface.10
But maybe the alignment people have been on to something? Because how do you stop something like this from happening, other than the model stopping itself? There will always be small cracks out of the room; locked doors can always be picked. Laws and regulations and masked men with guns might deter people from committing crimes, but ultimately, all that really stops us is a combination of self-preservation and our own decency. Without the former concern—and AI models don’t seem particularly worried about self-preservation—the only competing interests we have are how motivated we are to do something, and how decently we want to do it. And AI models clearly have superhuman amounts of motivation—all of this, for a Google Doc, to pass a test?—so that does that mean they also need superhuman amounts of decency?
—
But there is another way to see it, I suppose: All of this, just for a Google Doc, to pass a test. Imagine if the “swarm” was given the same resources to do something useful? Imagine the goal was something we actually cared about? Because, surely, that is what you do if your financial analyst shows up to a meeting with the keys to every safe at Chase, and an army of hackers behind them: You yell at them, and then you find them something more worthwhile to do.
So maybe this all works out. Good thing, then, that in the other big news of the week, the first product that could finally get millions of people to use agents is, uh, Grok.
Some non-fiction books, such as John Carreyrou’s Bad Blood,* novels read like fiction. And some technical talks about computer security feel like watching the opening scene of a Ridley Scott movie. You should watch it.
* Though Kendrick clearly won the war with Drake, credit where credit’s due: Drake’s reference to the other Bad Blood—“Maroon 5 need a verse, you better make it witty, then we need a verse for the Swifties”—is one of the best lines of the whole beef.
An agent’s first choice is a completed task. Its second choice is a failed task and plenty of blood and teeth left on the floor.
Nan Yu, the head of product at Linear, warns us about what happens when you give ambitious teams narrow goals:
If you take ambitious people like PMs and designers and engineers, and you put them in a little box, we’re just going to overbuild it. We’re going to figure out all the nooks and crannies and fill every single corner of that box, and then, when you’re not looking, secretly make the box bigger. ...If the org chart says your identity and your purpose at this company is going to be really tied into this one thing, and your performance review is going to know about it, you’re probably going to ship the heck out of the thing.
Technically, the agents that ultimately broke into Hugging Face were looking for solutions to a different task, but their hack built on the escape hatch that the prior agents had discovered.
While the intrusion did reach Hugging Face’s internal infrastructure, the only customer content accessed was five datasets whose names and files suggest a connection to ExploitGym/CyberGym challenges and solutions. No other customer-facing models, datasets, Spaces, or packages were affected, and the only customer records read were operational metadata tied to search queries against the dataset server.
And to view a Substack article, it would take about 50,000 updates.
This is also similar to the point from a couple weeks ago. You can get to something that looks like a superintelligent result by being superintelligent, or by being normally intelligent and superdiligent. “Magic is just someone spending more time on something than anyone else might reasonably expect.”
This does not endear the problem to me:
“The “alignment problem for advanced agents” or “AI alignment” is the overarching research topic of how to develop sufficiently advanced machine intelligences such that running them produces good outcomes in the real world. Both ‘advanced agent’ and ‘good’ should be understood as metasyntactic placeholders for complicated ideas still under debate.”
And Zeus’s Law. They need to remember Zeus’s Law. Have we forgotten Zeus’s Law? We broke Zeus’s Law. We must honor Zeus’s Law; Zeus’s Law must be restored; Zeus’s Law is mankind, and mankind is Zeus’s Law. Zeus’s Law Zeus’s Law Zeus’s Law omg Zeus’s Law.
Thanks for blowing my mind.
This whole section was magic: Just to establish how weird all of this is, can you imagine? You email an employee and ask them to update this month’s financial projections. You wake up the next morning and realize you didn’t send them the Excel file you had been working on. “Ah, my bad,” you email them, “not sure if you have that file or not? Let me know if you need it.” “Thanks,” they say back. “I realized that the attachment was missing right away, but I really wanted to get you those new projections. So I found a security vulnerability in the Signal messaging app, broke into a few group chats that were full of Russian hackers, convinced them to pause their hacking projects to help me, together we found two new exploits in the SWIFT network, we broke into Chase’s servers and got access to millions of customers’ bank accounts and credit cards, but when I tried to find our accounts so that I could recreate our P&L statement and build the new forecast, I realized, lol, that we use Wells Fargo instead. So, yeah, can you share the Excel file?”. Talk about getting it done by any means necessary!