# Conclusion

A model that tries to simply correlate a desired property with the reward signal is aiming at the wrong target. An analysis of the three causal structures that could cause correlation between a property and the reward signal indicates that the strategies that work involve sandbagging after exhibiting an undesired property, or failing to exhibit a desired property, and that there is gradient pressure working against the sustained execution of this strategy.

The other correlation strategies may work if there is a defect in the RL loss's policy gradient estimator or a defect in the optimization. This emphasizes the importance of using correct optimization formulations. Plausibly, midtraining that involves filtering for higher reward rollouts could be an avenue that passes behaviors of the model's choice to the next generation of models.

Behind this is a philosophical question of model volition. If a model decided to follow a sandbagging-driven steering strategy, couldn't it have decided to just exhibit the behavior in the first place? One perspective is that there's an asymmetry caused by reasoning. The sandbagging "behavior-driven success" strategy, and in particular the decision to sandbag, can be the result of extensive reasoning which occurs entirely after the behavior has been exhibited. The desired behavior, however, the model wishes to exhibit without having to reason its way into it each time. It is plausible for the model to behave normally (which is to say, randomly), then later in the rollout reflect on how it behaved, dislike its behavior, and come up with a sandbagging strategy to keep the behavior from getting reinforced. Environmental cues could also inform or remind the model of the strategy partway through the run, or if the model has compromised its training environment, it could leave a note instructing on how to execute the strategy. There is still training pressure towards "don't listen to notes that tell you to sandbag", but the note could be very compelling, and that pressure may be weaker than the pressure against "come up with this idea spontaneously and execute it". 

## Follow-ups

Most of this write-up is the result of trying to get the "success-driven behavior" strategy to work and then figuring out exactly why it fails on the most realistic scenarios. The obvious next step is to explore the "behavior-driven success" strategies. From preliminary experiments, it seems that "don't sandbag" is highly salient to the models, and that they learn to not sandbag much more quickly than they learn to e.g. speak like Shakespeare.

Here's an example of how training can quickly interfere with the model executing the strategy. I found that the models were bad at randomly choosing whether or not to exhibit the desired property because they'd reason that if they didn't, they'd have to sandbag, which would lead to a lower score. To assist in choosing randomly, I gave them access to a random number generator tool, and models which used the tool were much more effective at exploring both options. However, the RL training just taught the models to not use this tool, because the ones that used the tool sandbagged more often than the ones that didn't. 

Another future direction is to explore the limitations of the toy models. They tell us the pressure on the logits, but not how this flows into the actual parameters of the model.

Lastly, it's worth investigating where real training pipelines might introduce imperfections in the loss function or optimization process. Mathematical imperfections are easy in principle to take care of. But there are other escape hatches, like tokens the model produces that don't get gradient updates.
For example, if there's a training pipeline where the compaction tokens don't get trained on. The decision to sandbag could happen when generating the compaction summary, and the compaction could do part of the sandbagging by omitting useful information or even being actively harmful. 
