Veil

Research

In modern language models, post-training appears to induce behavioral structures that generalize across tasks, contexts and capabilities [1]. At the same time, recent interpretability work suggests that models develop globally accessible representational workspaces [2], while recent training research shows that behavioral traits can transfer through data that appear unrelated to those traits [3].

Our current goal, which we expect to change as the field evolves, is to develop a better understanding of whether internal objects produced by training can be made sufficiently stable, interpretable and aligned to safely participate in training and supervision of increasingly capable successor systems.

As our work develops, we expect to organize this agenda into more specific research directions.

Reading

  1. [1]

    Marks, S., Lindsey, J., & Olah, C. (2026, February 23). The persona selection model: Why AI assistants might behave like humans. Alignment Science Blog. https://alignment.anthropic.com/2026/psm/

  2. [2]

    Gurnee, W., Sofroniew, N., Pearce, A., Piotrowski, M., Kauvar, I., Chen, R., Soligo, A., Bogdan, P., Ong, E., Wang, R., Thompson, T. B., Abrahams, D., Kantamneni, S., Ameisen, E., Batson, J., & Lindsey, J. (2026, July 6). Verbalizable representations form a global workspace in language models. Transformer Circuits Thread. https://transformer-circuits.pub/2026/workspace/index.html

  3. [3]

    Cloud, A., Le, M., Chua, J., Betley, J., Sztyber-Betley, A., Hilton, J., Marks, S., & Evans, O. (2025, July 22). Subliminal learning: Language models transmit behavioral traits via hidden signals in data. Alignment Science Blog. https://alignment.anthropic.com/2025/subliminal-learning/