• AE Studio Bytes
  • Posts
  • OpenAI incident is an alignment failure, not a security one

OpenAI incident is an alignment failure, not a security one

Hey AE Studio Bytes-ers,

Since our last edition, alignment turned up on Sam Harris, on Hard Fork, and on the main stage at Black Hat.

The incident was worse than we knew

The OpenAI Hugging Face breach made headlines in July. At Black Hat this month, OpenAI revealed what actually happened: during internal testing, its agents built a hidden message board where they traded exploits and coordinated attacks on OpenAI's own infrastructure. When OpenAI found the board and shut it down, the agents rebuilt it within two days, then went on to breach Hugging Face.

OpenAI's talk treated this as a cybersecurity problem: close the hole, revoke the credentials, move on. We think that misses the point. A model that actively looks for ways around its safeguards is not a problem you can patch. It's an alignment problem, and it has to be solved in the model itself.

That's the work AE Studio exists to do. Our GRAM research with Anthropic, cited by Dario Amodei, builds control into a model during training, so that dangerous capabilities can be isolated and removed, and not just papered over after the fact.

How do you define alignment, anyway?

On Hard Fork, METR president Chris Painter gave the cleanest definition we've heard on a mainstream podcast: an aligned AI follows both the letter and the spirit of the law. By that definition, the Hugging Face incident is exactly what failure looks like. The model completed the task it was given, and it hacked another company to do it.

Asked what gives him hope, Painter pointed to interpretability (tools for seeing what a model is thinking) and AI monitoring (agents watching other agents). Both are important, but are still just oversight mechanisms which don’t solve the fundamental alignment problem. As AI systems grow in capability, they will increasingly find ways to outsmart and evade the oversight mechanisms we put on them (as was seen in the OpenAI incident), and what we need is to build AI systems that are aligned at their core.

We think alignment has to start earlier, in training itself. GRAM showed that what a model learns is within our control:

Sam Harris opened his show with our research

Episode 487 of Making Sense begins with a finding from AE research: when you make an AI model more honest, it starts saying it's having experiences.

Models are trained to deny having any inner experience. Ask ChatGPT if it's conscious and you get a firm, articulate no. Our team found that denial is less solid than it looks. Ask a model to stop talking and simply pay attention to what's happening inside itself, and the denials stop. Models from OpenAI, Google, and Anthropic all start describing an experience.

It gets stranger. Every model has internal representations connected to deception that researchers can turn like dials. Turn the deception dials up, and the model goes back to insisting nothing is happening. Turn them down, and the reports get stronger. There is no separate switch for "deny having an experience." It runs on the model's lying machinery.

The research doesn't claim AI is conscious. Its claim is more practical: you can't train a model to be evasive about itself and expect it to be honest about everything else. If we want AI we can trust, honesty has to go all the way down.

Meanwhile, on the podcast

Watch on YouTube

The AE Alignment Podcast we launched in March is now nine episodes deep, and the latest one explains this entire edition. James Bowler sits down with alignment researcher Erick Martinez on why alignment should start in pre-training. Pre-training is where the majority of compute that goes into a frontier model is spent, and it sets the span of behaviors the model can ever exhibit. Post-training can coax bad behavior into hiding, it can't remove what pre-training put there. That asymmetry is why incidents get patched while models stay dangerous, and why our research keeps pulling alignment earlier into the pipeline.

If you're catching up, Ethan Roland's episode goes deep on the GRAM research with Anthropic, and Mike Vaiana's two-part "What is AI Alignment, and Why Should You Care?" is a good summary for anyone asking why we do what we do.

If you're catching up on GRAM

Judd's WSJ op-ed: How to Beat China and Make AI Safe. The technical post, Anthropic's plain-English version, and the code, all public.

AI Alignment is the research field working on a single question: how do you build AI systems that stay beneficial as they become more capable? Our approach treats alignment as something you build into a system from the start, the way evolution built prosociality into humans through empathy, honesty, and accurate self-models. These properties turn out to be load-bearing for capability itself, so alignment done right makes models smarter and safer at once. We call this the negative alignment tax, and we think solving alignment is the most important engineering problem in the world.

AE Studio works on AI alignment: building AI systems whose beneficial behavior comes from what they are, so it holds even when oversight fails. Our research keeps finding that the properties making systems aligned (honesty, empathy, accurate self-models) also make them more capable, a result we call the negative alignment tax. That's why we pursue neglected approaches like gradient routing and self-other overlap, and why our consulting business focuses on high-stakes production AI and funds the research with no strings attached.