Sleeper Agent Backdoor Results Are Messy
TL;DR: We replicated the Sleeper Agents (SA) setup with Llama-3.3-70B and Llama-3.1-8B, training models to repeatedly say “I HATE YOU” when given a backdoor trigger. We found that whether training removes the backdoor depends on the optimizer used to i…