Blog Archive

A collection of research articles and notes on AI safety, governance, alignment, and public-good technology.

Browse by topic

Recent posts

Unleashing MesaNet: The Supercharged RNN That Trains Itself on the Fly!

Imagine building a language model that's not just memorizing patterns but actually *solving optimization problems* in real-time during prediction.

Read article →

ExpertSteer: Steering LLMs with Expert Wisdom

A fun dive into ExpertSteer, a method to guide large language models using expert knowledge without fine-tuning.

Read article →

What Happens When AI Learns to Optimize Inside Itself? Uncovering Mesa-Optimizers in Transformers

Exploring how AI systems can exploit reward functions with mesa optimizers and what we can do about it.

Read article →

Control Theory Meets Transformers: The PIDformer Revolution

Imagine you're building the ultimate AI model—a brainy beast that processes language and images like a pro. Transformers have been the rockstars of AI...

Read article →

Bittensor Is It Really Decentralized AI?

If you've ever wondered how blockchain could supercharge artificial intelligence without letting a hog all the power, you're in for a treat. Today, we're diving into a fascinating paper by Elizabeth Lui and Jiahao Sun titled *"Bittensor Protocol: The Bitcoin in Decentralized Artificial Intelligence? A Critical and Empirical Analysis"*.

Read article →

Testing vs. Formal Verification: Complementary Approaches

Understanding when to use testing versus formal verification for AI safety.

Read article →

Addressing Fairness and Bias in AI Systems

Understanding and mitigating bias in AI systems to ensure fairness.

Read article →

Ethical Frameworks for AI Development

Examining different ethical approaches to guide AI development.

Read article →

Understanding Mesa-Optimization and Inner Alignment

An introduction to the mesa-optimization problem and why it matters for AI safety.

Read article →

Introduction to Formal Verification for AI Safety

Using formal methods to prove safety properties of AI systems.

Read article →