AI Alignment
CoT Monitoring
- The Most Forbidden Technique is not always forbidden
- Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models with Dewi Gould, Francis Rhys Ward, Anders Woodruff, and many others
- 13 Arguments About a Transition to Neuralese AIs
- Hidden Reasoning in LLMs: A Taxonomy with Rohan Subramani and Shubhorup Biswas
- Extract-and-Evaluate Monitoring Can Significantly Enhance CoT Monitor Performance with Rohan Subramani and Shubhorup Biswas
- On Recent Results in LLM Latent Reasoning
Continual Learning
all with Rohan Subramani, Owen Terry, Achu Menon, Zhijing Jin, Francis Rhys Ward, and Seth Herd
- Perspectives on Continual Learning: Survey Results and Forecasts
- Angles of attack for continual learning safety
- How might continual learning affect safety and alignment?
- What's Continual Learning, and Why Might We Expect To See It In Advanced LLM Agents?
- Implications of Continual Learning for LLM Agents: Introduction
Character Training
Misc AI Safety Writing
- Paper Summaries 2
- How we spent our first two weeks as an independent AI safety research group with Rohan Subramani and Shubhorup Biswas
- Paper Summaries 1
- Can Reward Be the Optimization Target of an LLM?
- A Dialogue on Deceptive Alignment Risks
- Evaluating the Goal-Directedness of Language Models with Elizabeth Donoway and Marius Hobbhahn
- Exploring the Lottery Ticket Hypothesis
- Clarifying the confusion around inner alignment