- AE Studio Bytes
- Posts
- AE Studio Bytes — GRAM momentum edition
AE Studio Bytes — GRAM momentum edition

Hey AE Studio Bytes-ers,
Three weeks ago we told you about GRAM, our research with Anthropic. It has been a busy three weeks.
Open-weights AI models, models anyone can download, run, and build on, power much of today’s AI innovation. The problem: once a model’s weights are public, its safety guardrails can be removed.
This week, Anthropic CEO Dario Amodei published Anthropic’s position on open-weights models, and pointed to our recent research with Anthropic as one of the promising methods for making them safer.
The approach: train a model so that dangerous knowledge, like advanced virology or cyberattack techniques, is kept in separate “compartments” that can be removed before the model is released. Everything else the model can do stays intact. It’s a step toward AI that is both open and safe, without trading one for the other.
Judd took the argument to The Wall Street Journal. In his op-ed, “How to Beat China and Make AI Safe,” he makes GRAM the centerpiece of a national strategy: refusals are safety theater because the knowledge remains in the weights, while a GRAM-trained model behaved as if it never learned the dangerous information, and held even when attackers tried to train it back in. His argument: build the world’s most capable open-weight model in America with the dangerous capabilities stripped out before release, and make Beijing’s giveaways second-best.
Fortune’s Eye on AI featured the research. Their framing: want to make AI models safer? Switch off the dangerous bits. And their read of where it points: “a future where AI labs could sell one model with different capability tiers unlocked for different customers, instead of training and maintaining a whole fleet of separately restricted models.”
This was an ICML 2026 paper.
Paper and code are public: technical post · Anthropic’s plain-English version · code
—————————————————————————————————————————— |
AI Alignment is the research field working on a single question: how do you build AI systems that stay beneficial as they become more capable? Our approach treats alignment as something you build into a system from the start, the way evolution built prosociality into humans through empathy, honesty, and accurate self-models. These properties turn out to correspond with capability gains, so alignment done right makes models smarter and safer at once. We call this the negative alignment tax, and we think solving alignment is the most important engineering problem in the world. If you’re new to AI alignment, start with our introduction to AI safety and alignment here. |
AE Studio pursues neglected approaches like gradient routing and self-other overlap, and our consulting business focuses on high-stakes production AI and funds the research with no strings attached. |
/image