SEP 8, 2026 · PREPRINT
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
arXiv
This is a benchmarking study of AI agent capabilities in mechanistic interpretability research, establishing measurement methodology but not generating evidence about model safety, alignment, or real-world discovery outcomes.
Reported
Gemma Scope feature dictionary si…131K+
Number of agent configurations ev…10
Number of tasks evaluated20