Trained steering vectors may work as activation oracles
Inspired by @Eriskii’s recent finding that trained steering vectors can teach a base model to act as an assistant, I replaced the Activation Oracle paper’s trained LoRA with a far smaller set of per layer trained steering vectors and found surprisingly…