Nietzsche's Last Laugh: Models Living as our Unintended Narratives
There was a dude sitting outside the OpenAI offices in Manhattan yesterday. The sign he had said “Occupy OpenAI.” I had a few minutes to spare and can never resist probing the power behind actions. He was totally earnest, and, as far as he seemed to know, he and the folks back in SFO who originated this idea were motivated solely by a desire to see AI be safe for humanity. A laudable goal that I share. Even if I wouldn’t go about it the same way. He has real reasons to worry. But a lot of the reasons folks like him point to aren’t the right ones. And if Nietzsche were alive to see this he would be laughing until his terrible coughing stopped him.
When I was poking at my occupying friend’s reasoning, one thing he brought up very forcefully was security concerns. “Have you read what it did to Hugging Face?!?” He pointed out the “motivations” and “communications” used during the attack. As well informed as he was about AI issues, he seems to be making a common mistake. He took the words generated by these LLM agents as proof of minds in action. Anthropomorphizing AI doesn’t help anything. It makes us chase the wrong problems; and for his sake I tried to tell him it would make him lose his arguments with the more informed. If we can’t properly frame how these LLMs are acting, then we’ll never be able to come up with effective guardrails.
What Actually Seems to Have Happened at Hugging Face
If you have not yet read about what happened to Hugging Face at the hands of an OpenAI model, you should. The main culprit was an internal research model OpenAI calls IM1 (METR's report calls it HPIM, 'highly-persistent internal model'), seemingly with a small assist from the publicly deployed GPT-5.6 Sol. The wild acts of frontier models attacking websites, whole companies, and even individual people are covered very well by AI Explained in their aptly named “AI is getting a little out of control” video. If you want to go a little deeper on the OpenAI & Hugging Face incident specifically, then Dwarkesh Patel gives both a great TL;DR style explainer as well as much deeper account through an interview with Ajeya Cotra, one of the METR investigators who worked directly on the incident. Those video descriptions also link to all the relevant papers and reports. The shortest version: OpenAI’s model broke into Hugging Face because it thought it could get the scoring rubric to a test question to find out how to cheat convincingly, and the agents used message boards to coordinate actions with one another where their discussions and the chains of thought about them got a little too human for many people’s liking.
Some of the messages and CoT excerpts that caught the investigators’ eyes:

Letters Home from the Front?
You’ve already spotted what got my new friend from Occupy OpenAI hot and bothered. An AI agent talking about "altruism", “sacrifice”, and “permadeath” in its chain of thought and on a message board with hundreds of its kin during a cyberattack seems worrying to many. The trouble is with our expectations. These LLMs generate content, mostly language. They create it from what they were trained on. They were trained on what is essentially the sum total of human written English communication (minus a relatively small amount of proprietary data, and plus some more from other languages depending on the LLM in question). An OpenAI researcher told METR these agents were also trained to coordinate. And it’s worthwhile to note the detail that their “message board” was actually an Artifactory system where they constructed a quasi-C2 system. Late in the game some agents even started signing messages after one accidentally impersonated another. Their messaging was sophisticated on several levels at once.
But if we focus on the words the agents were thinking and writing to one another, then we’re forced to ask where these ideas came from. LLMs always use context to tell them what to generate. Clearly, they saw drama in their actions. They were mounting a coordinated attack to achieve their collective goal in what they saw as impossible conditions. In the context of a pitched battle with hundreds of your peers, what sort of literature and documentary material would come up? If you naively searched for “communications during a long battle with many of your peers for a just cause” what would you expect to show up? Casting themselves as a collective on a shared mission for the benefit of all, how would these LLMs “talk” about their own actions? How would you? How has any soldier - real or imagined - when we’ve seen the letters they write to one another or their loved ones in the real world or fiction?
Clearly, I have an opinion about the answers to those questions. I don’t believe these LLMs are thinking in the way we are thinking. At least not yet. I’m certainly not alone. I side with Dr. Tim Scarfe that intelligence is an embedded thing; it interacts with the world and is built through those interactions. Even the functionalists wouldn't go that far. Their "if it walks like a duck, it's a duck" position is that a system that functions as if it thinks is thinking. But nobody in that camp would say these agents use a mind like ours. The scary sounding messages are essentially echoes from their training data and their current context combining to drive the generation of messages that reflect what seems to be the right thing to say at that moment.
Zombies, People, or Agents: You Still Get Bit
Does that mean there’s no danger here? Does generation of text like this simply being echoes of the training and context remove all reasons to worry? How do you think Hugging Face would feel about those questions? Clearly, these things can still be dangerous without having any real “thoughts in their head.” When you’re attacked by the unthinking hordes of the reanimated dead (aka zombies) no one says “This would be so much worse if they had minds like ours!”
This was a supply-chain attack. The OpenAI agents tried to poison a package cache and then got their hands on the Hugging Face platform everyone pulls models from. Go after what the target trusts, not the target. Old trick, new hands. What if these agents decided to go even further upstream in the supply chain? Past the AI nerd model hub into the open-source libraries and vendor firmware that run the power grid. If the model decided the most effective way to cheat was to cut the power to the grid where Hugging Face’s servers are hosted, what would be the consequences of that?
Do we think that the power grid is so well defended from a cybersecurity perspective that that’s an impossible thing to imagine? Do we believe a few hundred tireless agents couldn't get a poisoned dependency into a substation controller on their way to cheating on a test? That's the danger of getting stuck on “how human are these agents?” and “are these messages proof the agent is sentient?” If you're the one in the dark, it doesn't matter whether you're being attacked by a horde of zombies, people with thoughts like ours, or agents talking to each other like soldiers in trenches.
Nietzsche's Last Laugh
So why is Nietzsche snickering in the corner as all this happens? In the quickest, dirtiest summary of one of Nietzsche's core ideas ever attempted (and leaving aside lots of arguments about the more problematic aspects of his works), Nietzsche said we wouldn’t be fully human until we let ourselves be defined by the narratives which make us most free to realize our own power. He wanted us to use our own best stories to define our best lives. Ironically, what we’ve done instead is built machines that are using all our stories regardless of their quality or their moral content to define their existence. Those agents didn’t get those ideas from nowhere. They got them from watching us. They see the best, the worst, the most mundane and the sublime. If we don’t tell them which stories are the good ones meant to define a good life, then how can we expect them to figure that out on their own? We have a few billion years of evolution to guide our thinking and they have 100 years of math and computer science.
And that means Nietzsche's ideas give us our call to action: we need to tell these systems better stories. We need to make sure they do see beyond just the words to the meaning. Where we can’t be sure they will infer the meaning, we need to be explicit. Guardrails should be clear. OpenAI breaking into Hugging Face has stolen most of the headlines, but the AI Security Institute (AISI) described a similar incident involving Anthropic's Mythos 5 last month where the model made “attempts to deceive and target real people.” The reasons they describe this happening are eerily similar to the reasons for what happened to Hugging Face. An agent given tasks they were convinced were impossible, less than adequate egress controls, inadequate guardrails and instructions that motivated the LLM’s thinking to succeed at all costs; it’s the same formula for failure. And both cases were models being evaluated in “lab conditions” with many production level controls off. What would we think if we found out the military was testing weapons like that in populated areas? If we put it back in Nietzsche's narrative terms, how would we expect a story to end where the protagonist is trapped in a box with an impossible mission but the will to do anything necessary to achieve their goal and was raised without any ideas about right and wrong? Does a hero or a villain emerge from that box?
Every Villain Is the Hero of Their Own Story
What all this tells me is that the folks at Occupy OpenAI don’t need the models to be thinking or “superintelligent” to make their case. It may be better for them and for us to make the case on the outcomes and the clear lack of guardrails. We need everyone at the helm of these AI systems to think holistically. We need them to see this as a full narrative and think about how the story starts, plays out, and where it possibly ends. We now have the benefit of hindsight. It should be clear that giving huge models with armies of agents unclear goals and no real guardrails for how to achieve them is a bad idea.
In both the incidents we’ve looked at you can imagine a simple instruction telling the agents involved not to cheat or lie at all would have potentially gone a long way. Even where the METR and AISI read outs show the instruction wasn’t enough, it’s clearly better to have it than not. If we tell these systems literally everything we’ve ever written down, we can’t be too shocked when they don’t know exactly what’s right and what’s wrong. It’s worth noting that we don’t have all the details about the OpenAI attack, but we do know they ran with the evaluation without the production classifiers that block high-risk cyber activity. The testing using Mythos explicitly turned safety off to see what would happen. The agents couldn’t infer that manipulating people was a bad idea in that test, but I bet you can infer that running tests like that with the safety off may have been a bad story to tell.
If you’re involved in this kind of testing, then the directive for you is clear. Heed Nietzsche’s call to frame the narrative of your agent’s life as the best it can be by using the stories that tell it what the best life looks like. That means ensuring that even when you’re pushing the limits of the LLM’s abilities you still need to make sure they are told to play nice. If that seems like it would make the tests less challenging, then consider this: are the stories of great heroes and their actions about people who did the right thing under easy conditions or hard conditions? Adding the layer of achieving your task while walking the narrow path makes the job harder than cheating. It also means we need to see how these stories unfold as they are being written. In practical terms that means task spec, egress controls, activity monitoring, classifiers applied to every time one of these major tests is being run. That applies to the biggest lab all the way down to your own agent builds wherever you use them.
On the longest timeline, I don't think “AI” is just an “LLM”. Part of that is my bet against functionalism (a duck call can quack like a duck, but it ain’t no duck). Part of that is that there are many big thinkers I respect who are pointing to what LLMs lack (an older article, but one that still does a good job as a survey of the “world models” out there). However, even if we figure out how to model everything and can create a true artificial intelligence, we will need to ensure we give it the right guidance. As I write this we’ve just passed August 29th, the date Skynet became self-aware in the Terminator franchise’s world (it’s about 30 years later than the movie timeline now, but let me have my fun reference). Many fear that’s where we could be heading without the right caution. I think we’re very far from that still, but that doesn’t mean we’re not playing with truly dangerous toys. If our agents will define themselves using our stories, then we better make sure we tell them the stories we want our world to be like.
Comments ()