Steering Might Stop Working Soon
Steering LLMs with single-vector methods might break down soon, and by soon I mean soon enough that if you’re working on steering, you should start planning for it failing now.This is particularly important for things like steering as a mitigation agai…