# Steering is hard.

A worry is that a model might be able to use situational awareness and knowledge of training to "steer" itself during RL. When training such a model, alignment becomes much more difficult because the model developer needs to ensure the model is aligned over the entire training trajectory, not just at the final checkpoint.

**Hypothesis**: A model can steer its RL training by deliberately correlating the property it wishes to exhibit with the reward signal.

**Finding**: Correlation is not sufficient. Experiments showed success only when the loss function deviated from the correct policy gradient, pointing at an incomplete understanding of training dynamics. The main outcome of my investigation is a toy model that shows why a naive strategy will fail and which strategies might work, and an understanding that just being correlated with the reward signal is not sufficient for policy gradient methods to reinforce a behavior. A more complete analysis of policy gradients narrows the space of strategies, and in a toy model, the strategies that remain are self-defeating, in that engaging in the strategy reduces the model’s ability to continue engaging in the strategy in the future.

This page presents what I consider to be the more interesting parts of the research arc, including selected experiments, some general RL loss math that shows what gets learned, and its application to a toy model.

This was a solo project. All prose on this page was written by me; formatting and typesetting by Claude. All errors are mine.