Towards an understanding of safe personas

Our Research
fig. 1 — a stable configuration

Veil is a nonprofit research organization working on alignment science, interpretability, and conceptual research.

Within the decade, AI systems will plausibly match or exceed human researchers at improving AI itself. What comes out of that loop depends on the character of the models running it — and a model that behaves well under observation tells us little about the one left unobserved. We study what personas are made of, and what would justify trusting one.

Research directions

[1]

Behavioral dynamics

A trained model holds many stable ways of behaving, and a small push — a prompt, a persona, a fine-tune — can carry it from one to another. We study these transitions directly: why jailbreaks work, when they resemble phase changes, and what makes the assistant basin deep or shallow.

[2]

Interpretability

Behavioral evidence only goes so far; we want the mechanism. We look inside models for the structures that carry personas and dispositions, aiming at explanations precise enough to predict when behavior will hold and when it will give way.

[3]

Conceptual research

The questions that come before experiments: what alignment actually asks of a system, what would count as evidence for it, and which of today's framings will still make sense as models get stranger.

fig. 2 — a stable configuration

Team

We're a small group and we like collaborators. If our directions overlap with yours, write to us.

Blog

Notes, negative results, and write-ups as the work matures. First posts arriving fall 2026.

About

Veil was formed by Thomas Jou and David Crispell in July 2026 to investigate agent modeling, stability, and generalization.

We meet weekly to read, argue, and run experiments. If you'd like to sit in on a meeting or hear when results go up, reach us at hello@veil.lc.