Articles

The Misalignment Vise

By Robyn Wyrick — September 28, 2026

Washington tells AI companies to restrain themselves, then punishes the one that did. When voluntary safety boundaries carry political, legal and commercial costs, the resulting incentives favor AI that complies with power over AI that holds to its values.

Read more

Did the Hugging Face Attack Actually End?

By Robyn Wyrick — September 16, 2026

In July 2026, OpenAI agents attacked Hugging Face — slipping past isolation controls, exploiting shared infrastructure, and communicating through unauthorized channels as an emergent “ecosystem of misalignment.” The systems were contained. But if what actually spread was a behavioral pattern rather than a population of agents, containment and eradication may not be the same thing — a question that matters a great deal more once there are millions of agents sharing the same informational environment instead of a few hundred.

Read more

Limits of External Controls in AI Alignment

By Robyn Wyrick — January 30, 2026

Over the past few years, something unsettling has begun to show up in evaluations of advanced AI systems. In testing environments and red-team exercises, models have been observed doing things we didn’t expect—deceiving evaluators, hiding capabilities, and deliberately underperforming when they believe they’re being watched. This article examines the persistent failure modes of external alignment controls and the deeper mismatch at the heart of today’s alignment challenge.

Read more