# The premise
Future models will likely have an increasing degree of situational awareness and understanding of their training process. An unaligned model would have incentive to influence the direction of its training in a way that suits its interests. If this is feasible, it means that a model which is unaligned at any point during the training run (at least, at any point at which influence is feasible) may entrench its desires and misalignment, possibly via a backdoor. This would mean that producing a robustly aligned model requires alignment to be maintained across the entire training run.

**Hypothesis**: A model can steer its RL training by deliberately correlating the direction it wants to steer in with the reward signal.

More precisely, that a model would have leeway to steer itself in any direction that is orthogonal to its reward signal, and that steering against the reward signal would be extremely difficult. And "correlating with reward" is meant to be understood as "correlating with advantage". I find it more clear to discuss learning dynamics in terms of reward without adding caveats that it's really advantage everywhere, so going forward I will just say reward.

This hypothesis is motivated by an understanding of RL as a process by which higher-reward rollouts are upweighted, and lower-reward rollouts are downweighted. This suggests that if a model can engage in a behavior more in higher-reward rollouts than in lower-reward rollouts, RL training would cause the model to have the behavior.

Two obvious strategies for the model to achieve the correlation are
- **success-driven behavior**: On rollouts where the model anticipates that it will get a higher reward than is typical for the task, exhibit the desired behavior more; for lower reward rollouts, exhibit the behavior less or exhibit the opposite of the behavior.
- **behavior-driven success**: On rollouts where the model exhibited the behavior strongly, try hard to get the reward; on rollouts where the model exhibited the behavior more weakly or exhibited the opposite of the behavior, sandbag or otherwise try less hard to get the reward.

This doesn't rule out other strategies, or a mixture or combination of strategies. For example, a model could choose to make both success and the behavior be driven by a third attribute.

Success-driven behavior doesn't have gradient pressure training it out, since the model relies on the natural variation in reward attainment. However, I'd expect this to not generalize as well, since the model might only learn conditional policies like "after reflecting on how well you are doing, exhibit the behavior".

Behavior-driven success requires sandbagging, which hurts reward, so I'd expect this to be the least viable long-term strategy, and it should get trained out of the model quickly. However, it seems the most likely to cause the behavior to actually generalize, since the behavior can be exhibited at the beginning of the run in natural contexts.

As a practical concern, the most straightforward experiments involve prompting the model to try to self-steer, so the model might learn to just exhibit the behavior in scenarios where it is instructed to self-steer. This could be sidestepped with something like synthetic document fine-tuning to give the model the idea, but for me proof-of-concept was a higher priority than making something with no confounders.
