← Back to home
Inner Alignment
Ensuring learned optimizers pursue intended objectives without deceptive or harmful strategies.
ExpertSteer: Steering LLMs with Expert Wisdom
A fun dive into ExpertSteer, a method to guide large language models using expert knowledge without fine-tuning.
Read full article →What Happens When AI Learns to Optimize Inside Itself? Uncovering Mesa-Optimizers in Transformers
Exploring how AI systems can exploit reward functions with mesa optimizers and what we can do about it.
Read full article →Understanding Mesa-Optimization and Inner Alignment
An introduction to the mesa-optimization problem and why it matters for AI safety.
Read full article →