[1]
Behavioral dynamics
A trained model holds many stable ways of behaving, and a small push — a prompt, a persona, a fine-tune — can carry it from one to another. We study these transitions directly: why jailbreaks work, when they resemble phase changes, and what makes the assistant basin deep or shallow.
[2]
Interpretability
Behavioral evidence only goes so far; we want the mechanism. We look inside models for the structures that carry personas and dispositions, aiming at explanations precise enough to predict when behavior will hold and when it will give way.
[3]
Conceptual research
The questions that come before experiments: what alignment actually asks of a system, what would count as evidence for it, and which of today's framings will still make sense as models get stranger.