Datasets and viewer resources for studying evaluation awareness, instrumental behaviour, sabotage, and monitoring in AI agents.
AI & ML interests
AI Safety, AI alignment
Recent Activity
View all activity
Organization Card
AI Safety and Alignment Group
AI safety · alignment · evaluation
The AI Safety and Alignment Group is based at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems. This Hub hosts our public datasets and trajectory releases for work on AI-agent evaluation, safety, and alignment.
Research projects
Browse all GitHub repositories, follow the group on Substack, or visit the group page.
models 0
None public yet