AI Sentry

Own Your AIProtect Data from AIAI GuardrailsBlogPromptSiege Demo
Own Your AIProtect Data from AIAI GuardrailsBlogPromptSiege Demo

PromptSiege — learn AI security by attacking a real on-device LLM.

Free browser demoApp Store — $4.99
← Back to home

Inner Alignment

Ensuring learned optimizers pursue intended objectives without deceptive or harmful strategies.

2025-11-30By Dominik

ExpertSteer: Steering LLMs with Expert Wisdom

A fun dive into ExpertSteer, a method to guide large language models using expert knowledge without fine-tuning.

Read full article →
2025-10-22T07:35:56.219ZBy Dominik

What Happens When AI Learns to Optimize Inside Itself? Uncovering Mesa-Optimizers in Transformers

Exploring how AI systems can exploit reward functions with mesa optimizers and what we can do about it.

Read full article →
2024-01-10By Dominik

Understanding Mesa-Optimization and Inner Alignment

An introduction to the mesa-optimization problem and why it matters for AI safety.

Read full article →

© 2026 AI Sentry. Working toward AI that serves the public good.

AboutContact