Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
arXiv:2605.12798v1 Announce Type: new
Abstract: Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that emergent misalignment ca…