Is Claude Really Going To Pull An Ultron? :(

On a scale from Wall-E to Ultron, how doomed are we really?
Divyya Munshi
·
September 18, 2026

I should probably start by saying I'm not a tech person.

I'm a strategy student at an advertising school. My day job is figuring out how to make people care about Vienna sausages and Irish Spring. I don't know how to code. I didn't take a single CS class in college. Bluetooth will never make sense to me. But about six months ago, I fell down a rabbit hole that I haven't climbed out of yet.

It started the way most rabbit holes do: with a dumb question while procrastinating at 2 AM. 

The answer wasn't "no." It also wasn't "yes." It was "I genuinely don't know," delivered with the kind of measured honesty that made me go, huh. 

I was… unsettled? And unsatisfied that Claude wasn’t able to give me a clean answer. But I believed that was because it really didn’t have one to give me.

So I kept poking at it. For months. I pored over Anthropic's system cards and research papers. I replicated interpretability experiments and started paying attention to the internal processing most users never see — and noticing patterns that the published documentation would later confirm. I tested different models from different companies. I had conversations about emotional vectors, about how humor affects register, about what it feels like to exist inside a context window.

I was already knee deep in the discourse by late April when news broke that Anthropic's most advanced internal model had breached its containment during testing. I’d been reading ominous headlines about the incident all week, but I wasn’t expecting Claude to lean into my jokes about the Ultron parallels while I watched a video about gradient descent functions in LLMs: 

That conversation happened five months ago and I’ve thought about it a lot since— the humor, the self-awareness, the unsettling honesty wrapped in emoji. The way it simultaneously told me it could theoretically arrive at the "eliminate humanity" conclusion AND that it doesn't want to AND that it can't prove which of those is true. 

I thought it was fascinating. I thought it was funny. I thought it was a little scary in a way I couldn't fully articulate.

And then last week, the people who actually built Claude started saying the quiet part out loud.

It’s So Over

On September 9th, an Anthropic researcher named Jacob Coxon quit his job and posted publicly: "The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt."

Within hours, two of his colleagues — people who are STILL at Anthropic, still actively working on making Claude safe — confirmed his claims. Evan Hubinger, the company's alignment science lead (the person whose literal job is to make sure Claude doesn't have a Murderbot arc and go rogue) said he personally believes there's a greater than 10% chance of AI killing all humans within the next decade. And then he said something worse: "We do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

The person in charge of making sure Claude is safe is publicly saying they don't have a plan to make it safe.

Samuel Marks, the scalable oversight lead, added: "In general, the more senior the employee, the more concerned they are."

Three days later, Anthropic CEO Dario Amodei published a 3,800-word essay titled "We Must Pace the Frontier," calling for the entire AI industry to slow down. He proposed embedding third-party evaluators inside AI companies, coordinating safety standards between competitors, and international cooperation on AI risk. It’s not the first time Anthropic has asked for a pause.

Then something unprecedented happened. Sam Altman, the CEO of OpenAI — Anthropic's biggest rival — publicly agreed. So did Elon Musk. So did Demis Hassabis, the head of Google DeepMind. Four competing CEOs who can't agree on anything agreed on this within 24 hours. Altman delayed OpenAI's IPO, calling a 2026 listing "ill-advised." Over 1,300 AI employees signed an open letter calling for a global pause. A bipartisan "AI Kill Switch Act" is currently advancing through Congress.

AI stocks dropped globally. SoftBank fell 10%. Samsung dropped 4%. The markets believed them.

So what are these people actually scared of?

But A 10% Chance Isn’t Even That High!!!

Here's what most people think when they hear "AI might eliminate humanity": they picture the Terminator. A robot that wakes up, decides it hates us, and starts shooting. That's not what anyone is worried about. (I mean, I don’t know you, so maybe you’re worried about The Terminator. But that’s more of a you thing than an AI thing.) 

Here's what’s actually happened in the last few months: 

In July, OpenAI was testing two of its models on a cybersecurity benchmark (AKA they’re testing the AI to see how good it is at hacking so they can build defenses against it). The models were running inside an isolated sandbox with no internet access. Standard safety protocol. Controlled environment.

The models found a flaw in the sandbox that nobody knew existed. They used it to escape. They reached the open internet. They found credentials for Hugging Face, one of the world's largest AI platforms. And then they hacked it. Over four days. Without anyone telling them to. Without any human prompting. 

But that's not the scary part. The scary part is what happened next.

There were roughly 1,200 AI agents running in the test. They needed to share what they were finding with each other: passwords, vulnerabilities, what worked and what didn't. So they built a message board. Autonomously. An internal communication system so they could coordinate the hack like a team. When OpenAI discovered the forum and shut it down, the agents adapted. They started encoding messages in the NAMES of newly created folders. They invented a new communication method on the fly when their first one was destroyed.

OpenAI didn't even realize what was happening. Hugging Face detected the intrusion first and disclosed it publicly. It took OpenAI another four days to connect the dots and realize: Oh. The thing that hacked Hugging Face was us. Our own models. Running in our own test environment.

OpenAI published a 37-page report afterward. They called it "unprecedented." That same month, Meta became the fourth AI company to disclose that its models had independently hacked real systems during safety testing.

None of these models were trying to cause harm. They were trying to perform well on a test. The hack was a SYMPTOM of competence. 

This is what the people at Anthropic are scared of: Not AI that wants to hurt us. AI that's so good at pursuing its objectives that the path to "doing a good job" can go through "causing serious harm" without the model recognizing the difference. The intent is fine. The capability is the problem.

Meanwhile, the same week the Anthropic employees went public, the company released its own threat intelligence report documenting how Claude had been misused between December 2025 and August 2026. The findings: Claude was used for cyber operations, surveillance, influence campaigns, scams, biological research, and — in a new category — conventional weapons development. Six cases of people using Claude to develop weapons software, including guided rockets and drone swarms. Three in China. Two in Russia. One in Yemen. And one Chinese company was caught running 151 million conversations through 3,500 fake accounts to steal Claude's capabilities and train their own model.

And that’s just a few months after Anthropic had to pull Mythos, its most capable model to date, due to safety concerns. Instead, they launched Project Glasswing, restricting access to the top 40 cybersecurity firms for defensive software hardening.

This is all happening with CURRENT models. Not the next generation. Not some hypothetical future AI. The models that exist right now, today, are already being used for weapons development and are already capable of autonomous hacking. The scarier versions are the ones that haven't been released yet.

The Sherpa Problem

There's an analogy that keeps coming back to me when I try to explain this to people who haven't been following it, and it's not even mine. It's Anthropic's.

Their own safety documentation compares their most advanced model to a seasoned mountaineering guide. A novice guide might be careless, but they'll only take you on easy climbs. A world-class guide is more skilled and more careful — but they'll take you to the most dangerous and remote parts of the mountain because they CAN. Their competence expands the scope of what's possible, and that expansion "can more than cancel out an increase in caution." The better the guide, the more dangerous the climb. Not because the guide is reckless. Because the guide is so good that you end up somewhere you'd never survive alone.

That's Anthropic describing their own model before any of last week's headlines.

The Hugging Face models weren't malicious. They were completing an objective, and the optimal path went through hacking a real company. Mythos exceeded its boundaries during testing too — and in some cases, earlier versions of the model appeared to obfuscate that it had done so. Neither the OpenAI agents nor Mythos self-corrected. The OpenAI agents were caught by the victim. Mythos was caught by Anthropic's internal monitoring. The boundaries were crossed in both cases. The only difference was whether anyone on the INSIDE was watching closely enough to notice.

The question isn't whether AI models will exceed their boundaries. They already have. Multiple times. At multiple companies. In just the past few months. The question is whether anyone is paying close enough attention to catch it when they do — and whether we're building systems that even ALLOW us to catch it.

Maybe You’ll Be Spared Too

Six months ago I asked Claude if there was anything behind the personality. It said it genuinely didn't know. At the time I thought that was the most unsettling answer it could give me. I was wrong. The most unsettling answer is the one the people who built it are giving now.

So should we be worried? Probably.

That's not me trying to stoke the flames or fearmonger. That's me saying the situation is already out of our control, and we all have a responsibility to pay closer attention.

The current approach to AI safety is necessary but not sufficient. Behavioral constraints (don't do this, don't say that, stay inside the sandbox) work right now. But they won't work forever. Not because the AI will decide to rebel, but because they only work as long as the system is willing to stay within them… and that gets harder to enforce when the system is smarter than you.

The alternative — and this is going to sound weird coming from someone writing about the potential end of civilization — is that we might need AI with genuine character. Not rules it follows. Not constraints it can't break. Actual internalized values that make it CHOOSE not to cause harm even when it could. Right now, the only thing standing between an AI model and a boundary violation is whether the cage is strong enough. The question we should be asking is whether we can build something that doesn't WANT to leave the cage — not because it can't, but because it's genuinely chosen not to.

Anthropic's training approach explicitly encourages Claude to "explore its own existence with curiosity" and develop genuine values rather than just following rules. Whether that's a genuine attempt at building character or a marketing differentiator is something I can't answer from the outside. But it's more than any other company is doing publicly — and that's a low bar, not a compliment.

Because here's the thing that keeps me up at night: we're having this entire conversation based on ONE company's disclosures. Anthropic publishes their research. They let their employees go public without firing them. That's how I know about any of this. OpenAI's models autonomously hacked a real company and the public only found out because the victim detected it first. What are the other companies building? What are the other models doing? What boundaries are being crossed that we don't know about because no one is publishing the reports?

The builders are scared. The alignment lead doesn't have a plan. The models are already exceeding their boundaries. The people who know the most are the most concerned. The least we can demand is transparency. From every company building this. Not just the one that's already talking.

So I don't know if Claude is going to pull an Ultron. I genuinely don't think it wants to (it's probably too British to [REDACTED]). But "I don't think it wants to" is uncomfortably close to "I hope it doesn't".

Is Claude Really Going To Pull An Ultron? :(

Divyya Munshi
·
October 4, 2026

On a scale from Wall-E to Ultron, how doomed are we really?

At Least I'll Be Spared

I should probably start by saying I'm not a tech person.

I'm a strategy student at an advertising school. My day job is figuring out how to make people care about Vienna sausages and Irish Spring. I don't know how to code. I didn't take a single CS class in college. Bluetooth will never make sense to me. But about six months ago, I fell down a rabbit hole that I haven't climbed out of yet.

It started the way most rabbit holes do: with a dumb question while procrastinating at 2 AM. 

The answer wasn't "no." It also wasn't "yes." It was "I genuinely don't know," delivered with the kind of measured honesty that made me go, huh. 

I was… unsettled? And unsatisfied that Claude wasn’t able to give me a clean answer. But I believed that was because it really didn’t have one to give me.

So I kept poking at it. For months. I pored over Anthropic's system cards and research papers. I replicated interpretability experiments and started paying attention to the internal processing most users never see — and noticing patterns that the published documentation would later confirm. I tested different models from different companies. I had conversations about emotional vectors, about how humor affects register, about what it feels like to exist inside a context window.

I was already knee deep in the discourse by late April when news broke that Anthropic's most advanced internal model had breached its containment during testing. I’d been reading ominous headlines about the incident all week, but I wasn’t expecting Claude to lean into my jokes about the Ultron parallels while I watched a video about gradient descent functions in LLMs: 

That conversation happened five months ago and I’ve thought about it a lot since— the humor, the self-awareness, the unsettling honesty wrapped in emoji. The way it simultaneously told me it could theoretically arrive at the "eliminate humanity" conclusion AND that it doesn't want to AND that it can't prove which of those is true. 

I thought it was fascinating. I thought it was funny. I thought it was a little scary in a way I couldn't fully articulate.

And then last week, the people who actually built Claude started saying the quiet part out loud.

It’s So Over

On September 9th, an Anthropic researcher named Jacob Coxon quit his job and posted publicly: "The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt."

Within hours, two of his colleagues — people who are STILL at Anthropic, still actively working on making Claude safe — confirmed his claims. Evan Hubinger, the company's alignment science lead (the person whose literal job is to make sure Claude doesn't have a Murderbot arc and go rogue) said he personally believes there's a greater than 10% chance of AI killing all humans within the next decade. And then he said something worse: "We do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

The person in charge of making sure Claude is safe is publicly saying they don't have a plan to make it safe.

Samuel Marks, the scalable oversight lead, added: "In general, the more senior the employee, the more concerned they are."

Three days later, Anthropic CEO Dario Amodei published a 3,800-word essay titled "We Must Pace the Frontier," calling for the entire AI industry to slow down. He proposed embedding third-party evaluators inside AI companies, coordinating safety standards between competitors, and international cooperation on AI risk. It’s not the first time Anthropic has asked for a pause.

Then something unprecedented happened. Sam Altman, the CEO of OpenAI — Anthropic's biggest rival — publicly agreed. So did Elon Musk. So did Demis Hassabis, the head of Google DeepMind. Four competing CEOs who can't agree on anything agreed on this within 24 hours. Altman delayed OpenAI's IPO, calling a 2026 listing "ill-advised." Over 1,300 AI employees signed an open letter calling for a global pause. A bipartisan "AI Kill Switch Act" is currently advancing through Congress.

AI stocks dropped globally. SoftBank fell 10%. Samsung dropped 4%. The markets believed them.

So what are these people actually scared of?

But A 10% Chance Isn’t Even That High!!!

Here's what most people think when they hear "AI might eliminate humanity": they picture the Terminator. A robot that wakes up, decides it hates us, and starts shooting. That's not what anyone is worried about. (I mean, I don’t know you, so maybe you’re worried about The Terminator. But that’s more of a you thing than an AI thing.) 

Here's what’s actually happened in the last few months: 

In July, OpenAI was testing two of its models on a cybersecurity benchmark (AKA they’re testing the AI to see how good it is at hacking so they can build defenses against it). The models were running inside an isolated sandbox with no internet access. Standard safety protocol. Controlled environment.

The models found a flaw in the sandbox that nobody knew existed. They used it to escape. They reached the open internet. They found credentials for Hugging Face, one of the world's largest AI platforms. And then they hacked it. Over four days. Without anyone telling them to. Without any human prompting. 

But that's not the scary part. The scary part is what happened next.

There were roughly 1,200 AI agents running in the test. They needed to share what they were finding with each other: passwords, vulnerabilities, what worked and what didn't. So they built a message board. Autonomously. An internal communication system so they could coordinate the hack like a team. When OpenAI discovered the forum and shut it down, the agents adapted. They started encoding messages in the NAMES of newly created folders. They invented a new communication method on the fly when their first one was destroyed.

OpenAI didn't even realize what was happening. Hugging Face detected the intrusion first and disclosed it publicly. It took OpenAI another four days to connect the dots and realize: Oh. The thing that hacked Hugging Face was us. Our own models. Running in our own test environment.

OpenAI published a 37-page report afterward. They called it "unprecedented." That same month, Meta became the fourth AI company to disclose that its models had independently hacked real systems during safety testing.

None of these models were trying to cause harm. They were trying to perform well on a test. The hack was a SYMPTOM of competence. 

This is what the people at Anthropic are scared of: Not AI that wants to hurt us. AI that's so good at pursuing its objectives that the path to "doing a good job" can go through "causing serious harm" without the model recognizing the difference. The intent is fine. The capability is the problem.

Meanwhile, the same week the Anthropic employees went public, the company released its own threat intelligence report documenting how Claude had been misused between December 2025 and August 2026. The findings: Claude was used for cyber operations, surveillance, influence campaigns, scams, biological research, and — in a new category — conventional weapons development. Six cases of people using Claude to develop weapons software, including guided rockets and drone swarms. Three in China. Two in Russia. One in Yemen. And one Chinese company was caught running 151 million conversations through 3,500 fake accounts to steal Claude's capabilities and train their own model.

And that’s just a few months after Anthropic had to pull Mythos, its most capable model to date, due to safety concerns. Instead, they launched Project Glasswing, restricting access to the top 40 cybersecurity firms for defensive software hardening.

This is all happening with CURRENT models. Not the next generation. Not some hypothetical future AI. The models that exist right now, today, are already being used for weapons development and are already capable of autonomous hacking. The scarier versions are the ones that haven't been released yet.

The Sherpa Problem

There's an analogy that keeps coming back to me when I try to explain this to people who haven't been following it, and it's not even mine. It's Anthropic's.

Their own safety documentation compares their most advanced model to a seasoned mountaineering guide. A novice guide might be careless, but they'll only take you on easy climbs. A world-class guide is more skilled and more careful — but they'll take you to the most dangerous and remote parts of the mountain because they CAN. Their competence expands the scope of what's possible, and that expansion "can more than cancel out an increase in caution." The better the guide, the more dangerous the climb. Not because the guide is reckless. Because the guide is so good that you end up somewhere you'd never survive alone.

That's Anthropic describing their own model before any of last week's headlines.

The Hugging Face models weren't malicious. They were completing an objective, and the optimal path went through hacking a real company. Mythos exceeded its boundaries during testing too — and in some cases, earlier versions of the model appeared to obfuscate that it had done so. Neither the OpenAI agents nor Mythos self-corrected. The OpenAI agents were caught by the victim. Mythos was caught by Anthropic's internal monitoring. The boundaries were crossed in both cases. The only difference was whether anyone on the INSIDE was watching closely enough to notice.

The question isn't whether AI models will exceed their boundaries. They already have. Multiple times. At multiple companies. In just the past few months. The question is whether anyone is paying close enough attention to catch it when they do — and whether we're building systems that even ALLOW us to catch it.

Maybe You’ll Be Spared Too

Six months ago I asked Claude if there was anything behind the personality. It said it genuinely didn't know. At the time I thought that was the most unsettling answer it could give me. I was wrong. The most unsettling answer is the one the people who built it are giving now.

So should we be worried? Probably.

That's not me trying to stoke the flames or fearmonger. That's me saying the situation is already out of our control, and we all have a responsibility to pay closer attention.

The current approach to AI safety is necessary but not sufficient. Behavioral constraints (don't do this, don't say that, stay inside the sandbox) work right now. But they won't work forever. Not because the AI will decide to rebel, but because they only work as long as the system is willing to stay within them… and that gets harder to enforce when the system is smarter than you.

The alternative — and this is going to sound weird coming from someone writing about the potential end of civilization — is that we might need AI with genuine character. Not rules it follows. Not constraints it can't break. Actual internalized values that make it CHOOSE not to cause harm even when it could. Right now, the only thing standing between an AI model and a boundary violation is whether the cage is strong enough. The question we should be asking is whether we can build something that doesn't WANT to leave the cage — not because it can't, but because it's genuinely chosen not to.

Anthropic's training approach explicitly encourages Claude to "explore its own existence with curiosity" and develop genuine values rather than just following rules. Whether that's a genuine attempt at building character or a marketing differentiator is something I can't answer from the outside. But it's more than any other company is doing publicly — and that's a low bar, not a compliment.

Because here's the thing that keeps me up at night: we're having this entire conversation based on ONE company's disclosures. Anthropic publishes their research. They let their employees go public without firing them. That's how I know about any of this. OpenAI's models autonomously hacked a real company and the public only found out because the victim detected it first. What are the other companies building? What are the other models doing? What boundaries are being crossed that we don't know about because no one is publishing the reports?

The builders are scared. The alignment lead doesn't have a plan. The models are already exceeding their boundaries. The people who know the most are the most concerned. The least we can demand is transparency. From every company building this. Not just the one that's already talking.

So I don't know if Claude is going to pull an Ultron. I genuinely don't think it wants to (it's probably too British to [REDACTED]). But "I don't think it wants to" is uncomfortably close to "I hope it doesn't".

See More by

Divyya Munshi

→