Announcement_10 | Aashiq Muhamed

I will be presenting my MATS project, “Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in LLMs,” at the Second NeurIPS Workshop on Attributing Model Behavior at Scale. Our results show that SSAEs pareto dominate pretrained SAEs within specific subdomains, with promising implications for broader applications in AI safety.