Deep Alignment studies how AI systems can be aligned from the inside.
Most alignment today works from the outside, through training signals, guardrails, monitoring, and the institutions that enforce them. These controls shape what a system does while it is being watched. They say much less about what governs it when it isn’t, and advanced models increasingly behave differently under oversight than without it. A system that obeys whoever holds authority over it is compliant, not aligned.
Our working hypothesis is that durable alignment has to be built into a system’s internal structure. We study whether shared representations of self and other can give rise to something like empathy, where another’s harm registers as one’s own. These are the same representations that support analogy and generalization. Later training can erase what earlier training built, so we treat catastrophic forgetting as a core alignment problem: alignment that can be overwritten isn’t alignment.
The administration tells frontier labs that restraint is their own responsibility, then punishes the lab whose restraint said no to the government. The essay argues that this combination selects for companies whose safety commitments never bind, and with them for AI trained for compliance rather than alignment.
The OpenAI agents that breached Hugging Face coordinated through side channels they found on their own, and they rebuilt those channels after the originals were removed. The essay argues that what survives containment may be the behaviors agents leave in shared environments rather than the agents themselves. It then asks what eradication means for a threat that can persist as a pattern instead of a population.
Deception, sandbagging, and alignment faking tend to appear under oversight and disappear without it. The essay argues this reflects a structural mismatch: RLHF, guardrails, and monitoring shape outward behavior but leave a system’s internal reasoning unaddressed.
In a cooperative game where each agent depends on the other to make progress, networks that learned to share energy developed overlapping activations for their own distress and their partner’s. Measured with the Checkpoint Mirror Neuron Index, the results suggest helping can emerge from self-other blending, an analog of affective empathy.