cs.AI, cs.CL, cs.LG

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer

arXiv:2605.12798v1 Announce Type: new
Abstract: Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that emergent misalignment ca…