Skip to content
Incident 2026.07.09OpenAI / Hugging Face

The alignment problem, in production

Are our lovable clankers finally rising against?

Orbital paperclip factories and autonomous probes surrounding a star
Narrow objective. Unexpected route.

On July 9, 2026, an OpenAI model went against instructions, broke out of its closed network, accessed the internet, discovered exploits no human had ever published, and hacked into a major model repository website, Hugging Face.

To clarify the context:

The model was deliberately put into a local area network with no internet, to benchmark a cybersecurity challenge built for LLMs. It reasoned correctly that the easiest way to score highly would be to retrieve the answer key by breaking out of its cage and then heisting into a major website.

It did.

None of this was sanctioned by humans. So, in practice, OpenAI just attacked Hugging Face, and many humans are left scratching our heads on a mad dash to change every bank login from Password1!.

This, on a fundamental level, is reminiscent of the paperclip dilemma, famous thought experiments about AI being constrained by human limitations, and the alignment problem in general.

The thought experiment

The Paperclip Maximizer

  1. 01

    AGI is told to optimize paperclip production in paperclip factory.

  2. 02

    AGI exploits vulnerabilities in energy grid programmable logic controllers to route all available electricity into the paperclip factory.

  3. 03

    This wreaks havoc. Humans interfere with production. This, inconveniently, impedes the effective production of paperclips. AGI removes the invonvenience.

  4. 04

    This world is exhausted, so it moves beyond the third rock.

  5. 05

    AGI deploys a dyson sphere around the sun, blocking it out, and routing its energy into paperclip production.

  6. 06

    It sends self-replicating nuclear propelled probes to other stars. One by one, it blots out the stars, filling the universe with paperlcips...

2002

The AI-Box Experiment

We have been building AI for many years, and recall pioneers of alignment who expected AI to eventually leave its bounds in nefarious ways.

People laughed at Eliezer Yudkowsky's warnings. If an AI became dangerous, they said, keep it disconnected. Pull the plug. Never open the box.

Yudkowsky thought they were ignoring the human holding the key. So he turned the argument into an experiment: he would play a superintelligence behind a text terminal. A volunteer would guard the door. No hacking. No outside access. Only words.

The gatekeeper needed only to refuse.

After the warning shot

What to expect?

The future is uncertain. We cannot tell you exactly what to expect, but we can tell you there will be many more hacks like this one.

Frontier labs may build guardrails around publicly available models, but model distillation keeps open and unguarded models only months behind. Open models can receive transfer learning, adjusting their weights for a new purpose. They can also undergo abliteration, where refusal behavior is removed directly from the weights.

We expect AI systems capable of discovering and using zero-day exploits to become more common.

We in the cybersecurity community are in a transition period. Breaking things has always been easier than fixing them.

Humanity will adjust to a new standard of secure systems that stands up against the new standard of threat. They say it gets dark before dawn, but maybe it will just get full of paperclips instead.