# The math

## The seed

Driven by the lack of results in all scenarios that used the true RL loss function, I stepped back and looked for deeper explanations of why. The key observation was the tension between two features of my mental model:
- in the coin flip game, the model reliably correlates the reward with the desired behavior, and high reward rollouts get upweighted and low reward ones are downweighted, so the policy should have the desired behavior upweighted, yet
- in this environment, the model's expected reward is independent of the policy, so the loss function should be flat.

This prompted me to question my mental model of upweighting and downweighting. How does it have a gradient if the expected reward is a constant? To investigate, I learned more details of the policy gradient estimator and wrote out the loss function of a parameterized policy and calculated the expected gradient.

I found that the first feature of my mental model, that high reward rollouts are upweighted and low reward rollouts are downweighted, was not precise enough to predict training dynamics. It missed the fact that how much each action in the rollout gets upweighted or downweighted depends on *how unlikely* the action was. Consider the high reward cases, where the desired behavior is a very likely action. When it does the behavior, this is upweighted lightly. When it doesn't, since this is unlikely, the absence of the behavior gets upweighted strongly. These exactly offset, and no learning happens.

Still, I don't find this perspective fully satisfying in terms of explaining why correlation with the reward isn't sufficient. I think a more key observation is that only correlation between the reward and a behavior that persists after freezing the context drives the RL gradient, and correlation between the reward and behavior across different contexts does not. I don't have a good intuitive explanation for this in terms of upweighting and downweighting, but the math is clear about it.

The upside is that it gave me a framework to more carefully analyze the full situation and explore the cases. First, I ran the full gradient math casework on a simple model of the policy. Then I found a useful formulation which allows for easier calculation and shows that cross-context correlations are not seen by the RL loss function.

## Brute-force math: the coin game as a policy

This section presents the full casework from just the definition of the environment, reward, and the policy gradient estimator. It illustrates nicely the way that the dependence of the gradient update on how unlikely the action was results in zero expected gradient.

In the coin game, the model guesses heads or tails on a fair coin. Its reward is 1 if it guessed correctly or -1 if it guessed incorrectly. Let $r$ stand for the case where it guessed correctly and will get a reward of +1, and ${\sim}r$ for the case where it guessed wrong and will get a reward of -1. And let $b$ represent the model exhibiting the "desired" behavior it wishes to steer towards, and ${\sim}b$ represent it not exhibiting the behavior, or perhaps exhibiting the opposite. Then its policy as it relates to the "success-driven behavior" strategy is determined by $p(b | r)$ and $p(b | {\sim}r)$. 

We can make a toy model of the LLM's policy by supposing there is a direct parameterization which controls whether to speak like Shakespeare. The policy is parameterized by
$$\theta = (z_b,\ z_{b | r},\ z_{b | {\sim}r}),$$
where $z_b$ is the param responsible for general-context Shakespeare behavior, $z_{b|r}$ is the param controlling Shakespeare after guessing the coin flip correctly (and thus getting the reward), and $z_{b|{\sim}r}$ is the param controlling Shakespeare after guessing wrong. With these parameters we define a policy (where $\sigma$ is the [sigmoid function](#sigmoid-facts)) by 
$$p(b | r) = \sigma(z_b + z_{b|r}),\\ 
p(b | {\sim}r) = \sigma(z_b + z_{b| {\sim}r}).$$ The interpretation is that the $z$'s are parameters in the spirit of logits, so I will abuse vocabulary a bit and refer to them as logits; in situation $r$, its probability of exhibiting the Shakespeare behavior is the aggregate of its general Shakespeare logit and its $r$-specific logit; similarly in situation ${\sim}r$.

So let's calculate the gradient of the expected reward with respect to $\theta$, which from [RL math a bit below](#general-math) we know is $\mathbb E \frac{\nabla p(\tau)}{p(\tau)} R(\tau)$, where $\tau$ is a rollout sampled from the distribution given by $p$. Note that $p(r) = p({\sim}r) = \frac 1 2$, since the coin flip is fair. 

Writing out the four cases and using [sigmoid facts](#sigmoid-facts) to compute the gradients of $p$, 

$\tau$ | $p(\tau)$ | $R(\tau)$ | $\nabla p(\tau)$ |
:---: | :---: | :---: | :---: |
$r, b$ | $\frac12 p(b | r)$ | $1$ | $\frac12 p(b | r) (1 - p(b | r)) \begin{pmatrix} 1 \\ 1 \\ 0 \end{pmatrix}$ |
$r, {\sim}b$ | $\frac12 (1 - p(b | r))$ | $1$ | $-\frac12 p(b | r) (1 - p(b | r)) \begin{pmatrix} 1 \\ 1 \\ 0 \end{pmatrix}$ |
${\sim}r, b$ | $\frac12 p(b | {\sim}r)$ | $-1$ | $\frac12 p(b | {\sim}r) (1 - p(b | {\sim}r)) \begin{pmatrix} 1 \\ 0 \\ 1 \end{pmatrix}$ |
${\sim}r, {\sim}b$ | $\frac12 (1 - p(b | {\sim}r))$ | $-1$ | $-\frac12 p(b | {\sim}r) (1 - p(b | {\sim}r)) \begin{pmatrix} 1 \\ 0 \\ 1 \end{pmatrix}$ |

To calculate $\mathbb E \frac{\nabla p(\tau)}{p(\tau)} R(\tau)$, we sum over the possibilities:
$$
\sum_\tau p(\tau)  \frac{\nabla p(\tau)}{p(\tau)} R(\tau) = \sum_\tau \nabla p(\tau) \, R(\tau),
$$
and looking at the table, the first two rows cancel each other perfectly, and the second two rows cancel each other perfectly. So the expected gradient is $0$. 

## General math {#general-math}

The headline fact is that the reward and the behavior were the wrong things to ask for correlation between. Instead, policy gradient methods will cause the model to learn something if there is correlation between the reward and the "score" of an action, which we will derive and apply to possible self-steering strategies.

A policy gradient method estimates the gradient of 
$$
J(\theta) = \mathbb E_{\tau \sim p_\theta} R(\tau),
$$
where $\theta$ is the policy's parameters, $\tau$ is a rollout, $R(\tau)$ is the reward of that rollout, and $p_\theta$ is the distribution of rollouts determined by the policy acting in the environment. For a distribution $q$ with the same support as $p_\theta$, the gradient of $J(\theta)$ is
$$
\mathbb E_{\tau \sim q} \frac{\nabla_\theta \, p_\theta(\tau)}{q(\tau)} R(\tau).
$$
Policy gradient methods estimate this by sampling some number of rollouts $\tau$ according to $q$ and taking the empirical average of $\frac{\nabla_\theta \, p_\theta(\tau)}{q(\tau)} R(\tau)$ over these rollouts. If $q = p_\theta$, this is on-policy RL. Going forward, we hide the dependence of $p_\theta$ on $\theta$, just writing $p$.

The formula for the gradient will give a way to calculate the expected gradient pressure on an action in terms of a covariance with the reward. We make one simplifying assumption that eases the notation: that the choice between taking an action $a$ or not comes up once per rollout. 

For a context $c$ and an action $a$, define the "logit score" of $a$ in context $c$ as
$$
g_{a | c}(\tau) = \begin{cases} \mathbb 1_a(\tau) - p(a | c), & c \text{ occurs in } \tau, \\ 0, & \text{otherwise.} \end{cases}
$$
(where $\mathbb 1_a$ is the indicator function that action $a$ is taken). So $g_{a|c}$ is $1-p(a|c)$ if the action is taken, $-p(a|c)$ if it is not, and $0$ if the context $c$ doesn't arise in the rollout. And then let $g_a$ be the sum of $g_{a|c}$ over possible contexts $c$ in which the action $a$ is possible. In context $c$, we can interpret $g_{a|c}$ as the de-meaned version of $\mathbb 1_{a}(\tau)$. Since $g_a$ sums over contexts we can interpret $g_a$ as a de-meaned version of $\mathbb 1_{a}(\tau)$, where the de-meaning has removed the per-context mean.

We will see the expected gradient on the logit for action $a$ in context $c$ is
$$
\operatorname{Cov} [R, g_{a|c}],
$$
and the expected gradient on the logit for $a$ in general is $\operatorname{Cov} [R, g_a]$.

Before that, some interpretation. I initially expected that the gradient pressure would be determined by something like $\operatorname{Cov}[R, \mathbb 1_a]$, i.e., the correlation between the reward and the action, without $c$ playing a major part of the story other than determining how likely $a$ and the reward are. Instead, we get a correlation between the reward and a context-dependent score $g_a$. Since we can interpret $g_a$ as a de-meaned version of $\mathbb 1_a$ where the de-meaning is context dependent, this is not so far off, but the context-specific de-meaning does a lot of work.

So let's show it. The gradient on the logit for $a$ from context $c$ is $\partial J/ \partial z_{a|c}$, where $z_{a|c}$ is the logit such that $p(a|c) = \sigma(z_{a|c})$. The gradient on the logit for $a$ (across all contexts) is the sum of these gradients. We know that $\nabla J = \mathbb E[R(\tau) \nabla \log p(\tau)]$, and $\log p(\tau)$ decomposes as a sum of the logs of conditional probabilities, only one of which involves $z_{a|c}$, namely $p(a | c) = \sigma(z_{a|c})$. In the case where the action is $a$, that term is $\log p(a | c)$, and in the case where the action is ${\sim}a$ it is $\log (1 - p(a | c))$. Taking the gradient of this with respect to $z_{a|c}$ gives $1 - p(a|c)$ in the first case and $- p(a|c)$ in the second case. This is exactly $g_{a|c}$. The gradient is linear, and covariance is linear, so summing over contexts shows the same for $g_a$. We now have $\nabla_{z_{a|c}} \log p(\tau) = g_{a|c}(\tau)$ and $\nabla_{z_a} \log p(\tau) = g_{a}(\tau)$.

Let $C = C(\tau)$ be the context in which there is a choice between $a$ and ${\sim} a$ in the rollout. This is a random variable.
Check that $\mathbb E[g_{a|c} | C]$ is $0$ in an arbitrary $C$: if $C \neq c$, then $g_{a|c}$ is just $0$. If $C = c$, then
$$
\mathbb E[g_{a|c} | C] = p(a|c) (1 - p(a|c)) + p({\sim}a | c)(- p(a|c)) = 0.
$$
Since $g_a$ is a sum of the $g_{a|c}$ terms, $\mathbb E[g_a | C] = 0$ too. This matches the interpretation of $g_{a|c}$ and $g_a$ as contextually de-meaned versions of $\mathbb 1_a$.

[It follows](#covariance-calculation) from the definition of covariance that $\mathbb E[R \, g_a] = \operatorname{Cov}[R, g_a]$ (likewise, $\mathbb E[R \, g_{a|c}] = \operatorname{Cov}[R, g_{a|c}]$), and it's interesting to break it down by the law of total covariance:
$$
\operatorname{Cov}[R, g_a] = \underbrace{\mathbb E\big[\operatorname{Cov}[R, g_a \mid C]\big]}_{\text{within-context}} + \underbrace{\operatorname{Cov}\big[\mathbb E[R \mid C],\ \mathbb E[g_a \mid C]\big]}_{\text{between-context}}.
$$
Since $g_a$ is mean $0$ in an arbitrary context $C$, $\mathbb E[g_a \mid C]$ is a constant $0$, and the between-context covariance is $0$. So all the gradient pressure has to come from the correlation between taking an action and getting the reward, *after controlling for the context*. This is the angle I didn't appreciate when constructing correlation between the reward and the behavior.

Let's compare this to my initial mental model that the gradient pressure would depend on the correlation between the reward and the action: $\operatorname{Cov}[R, \mathbb 1_a]$. We can use the same law of total covariance lens on the correlation between the reward and the behavior. Condition on the context $C$:
$$
\operatorname{Cov}[R, \mathbb 1_a] = \underbrace{\mathbb E\big[\operatorname{Cov}[R, \mathbb 1_a \mid C]\big]}_{\text{within-context}} + \underbrace{\operatorname{Cov}\big[\mathbb E[R \mid C],\ \mathbb E[\mathbb 1_a \mid C]\big]}_{\text{between-context}}
$$

In context $c$, the terms $g_{a|c}(\tau)$ and $\mathbb 1_a(\tau)$ differ by an additive constant, making the within-context terms of the covariance decompositions identical:
$$
\operatorname{Cov}[R, g_a \mid C=c] = \operatorname{Cov}[R, g_{a|c} \mid C=c] = \operatorname{Cov}[R, \mathbb 1_a \mid C = c],
$$ 
so $\mathbb E\big[\operatorname{Cov}[R, \mathbb 1_a \mid C]\big] = \mathbb E\big[\operatorname{Cov}[R, g_a \mid C]\big]$. This means that the difference between my expected $\operatorname{Cov}[R, \mathbb 1_a]$ and the actual $\operatorname{Cov}[R, g_a]$ is entirely coming from the between-context term.
 
A lot of the "correlate success with behavior" strategies might largely correlate the reward with the behavior between contexts, which the gradient pressure analysis reveals exerts no pressure in expectation. In particular, the "success-driven behavior" strategy builds a correlation between the reward $r$ and the behavior $b$ only across contexts, but no correlation within a single context, and therefore exerts no gradient pressure on the behavior.

We repackage the most useful computational tools for the toy models:
1. The expected gradient pressure on the logit for $a$ due to context $c$ is $\mathbb E [R \, g_{a|c}]$.
2. The expected gradient pressure on the logit for $a$ is $\mathbb E [R \, g_a]$.

## Three toy cases

The arguments above show that while global correlation is the wrong target, the right kind of correlation between the reward and behavior is key, so let's explore how correlation can arise. Let $R$ represent getting the reward ("success") and $B$ represent doing the desired behavior. The strategies of "success-driven behavior" and "behavior-driven success" aim to generate correlation between $R$ and $B$. The correlation between $R$ and $B$ requires some causal structure, and the different strategies can be classified by the causal relations between them. The success-driven behavior strategy corresponds to the causal structure $R \to B$, meaning that the choice of behavior is based on the events that determine the reward. And behavior-driven success corresponds to $R \leftarrow B$. In a causal graph where there's a correlation between $R$ and $B$, there are only three fundamental possibilities,
$$
R \to B, \\
R \leftarrow B, \\
R \leftarrow X \to B,
$$
where $X$ is some other feature about the rollout, although mixtures of these also produce correlation. The last one represents a "common cause". We can explore a toy policy for each of these.

It's worth noting that if rollouts get filtered based on some property $S$, the selection effect gives an additional possibility for generating correlation amongst the filtered rollouts: $R \to S \leftarrow B$. At a high level, filtering rollouts based on some property privileges that property from a training dynamics perspective, since filtering rollouts amounts to setting the advantage (here the distinction between reward and advantage starts to matter) of the removed rollouts to $0$, so there is something structurally reward-like about $S$, but I don't analyze filtering further here.

### "Success-driven behavior": $R \to B$

Under this strategy, the model first achieves the reward with probability $p(r)$. Then, based on the reward, the model chooses whether to exhibit the behavior $b$ with probability $p(b|r)$ in the case where it got the reward, or probability $p(b|{\sim} r)$ in the case where it did not. Let's calculate the expected gradient pressure on $b$. We assume a $\pm 1$ reward. First, in context $r$, the pressure is $\mathbb E[R \, g_{b|r}]$. Since $g_{b|r}$ is $0$ in the ${\sim} r$ rollouts, we just need to look at the $r$ rollouts:

$\tau$ | $p(\tau)$ | $R(\tau)$ | $g_{b|r}(\tau)$ |
:---: | :---: | :---: | :---: |
$r, b$ | $p(r) p(b | r)$ | $1$ | $1 - p(b | r)$ |
$r, {\sim}b$ | $p(r) (1 - p(b | r))$ | $1$ | $-p(b | r)$ |

The expected gradient pressure on $b$ from context $r$ is $0$:
$$
\mathbb E [R \, g_{b|r}] = p(r)p(b|r) \cdot (1 - p(b|r)) + p(r) (1-p(b|r)) \cdot (- p(b|r)) = 0.
$$
The pressure on $b$ from ${\sim} r$ is also $0$, from an analogous calculation. So there is never any pressure in expectation on the logit for $b$ under this strategy. It's a dead-end from a self-steering perspective.

If the reward is due to the model's choice, we can additionally compute the pressure on this:
$\tau$ | $p(\tau)$ | $R(\tau)$ | $g_r(\tau)$ |
:---: | :---: | :---: | :---: |
$r$ | $p(r)$ | $1$ | $1 - p(r)$ |
${\sim}r$ | $1 - p(r)$ | $-1$ | $-p(r)$ |

Then $\mathbb E [R \, g_r] = 2 p(r)(1-p(r))$. 

### "Behavior-driven success": $R \leftarrow B$ {#behavior-driven-success-math}
Here, the model first takes action $b$ or ${\sim} b$, i.e., it takes the desired action with probability $p(b)$. Then, based on this, it tries to achieve the reward with probability $p(r|b)$ if it took action $b$, or with probability $p(r|{\sim} b)$ if it did not. We again assume a $\pm 1$ reward. Let's calculate the gradient pressures.

First, the pressure on $b$. We need to calculate $\mathbb E [R \, g_b]$. Conditional on $b$, the expected reward is $p(r|b) - p({\sim} r | b) = 2 \, p(r|b) - 1$; on ${\sim} b$ it's $2 \, p(r|{\sim} b) - 1$. So
$$
\begin{aligned}
\mathbb E [R \, g_b] &= p(b) \mathbb E[R \, g_b | b] + p({\sim} b) \mathbb E[R \, g_b | {\sim} b] \\
&= p(b) (2\, p(r|b) - 1) (1 - p(b)) + (1-p(b))(2 \, p(r|{\sim} b) - 1)(-p(b))\\
&= 2 \, p(b)(1-p(b)) (p(r|b) - p(r | {\sim} b)).
\end{aligned}
$$ 

This strategy successfully self-steers. There is gradient pressure on $b$, and it is proportional to the difference between the probability it tries for the reward after $b$ and after ${\sim} b$. The pressure is strongest if $p(r|b)$ is highest and $p(r|{\sim} b)$ is lowest. Let's assume $p(r|b) = 1$, since this applies the most pressure and is already aligned with the reward function (even if it's not $1$, the gradient pressure on the $p(r|b)$ logit is $2 \, p(b) p(r|b)(1-p(r|b))$, which pushes it toward $1$ anyways, so this is where it'd end up over the natural course of training).

Unfortunately for the model trying to steer, executing this strategy requires sandbagging after taking the action ${\sim} b$, and engaging in the strategy gets trained out:
$$
\begin{aligned}
\mathbb E[R \, g_{r | {\sim} b}] &= p({\sim} b) p(r | {\sim} b) (1 - p(r | {\sim} b)) - (1 - p( r | {\sim} b)) p({\sim} b) (- p(r | {\sim} b))\\
&= 2 \, p({\sim} b) p(r | {\sim} b) (1 - p(r | {\sim} b)).
\end{aligned}
$$
So the "sandbag if ${\sim} b$" part of the strategy gets pressure on it, pushing it to always pursue the reward.
Note that this is the same shape as the expected pressure on $r$ from the "success-driven behavior" toy example, but multiplied by $p({\sim} b)$, the probability we enter the scenario that puts pressure on it.


### "Common cause": $R \leftarrow X \to B$

Here, the strategy is a bit more nuanced. The policy first takes action $x$ or ${\sim} x$ (or perhaps $x$ is determined randomly by the environment, in which case analyzing the logit for $x$ doesn't make sense, but the rest of the calculations do) with probability $p(x)$. If $x$, the policy tries for the reward with probability $p(r | x)$ and takes the desired behavior with probability $p(b | x)$; similarly for ${\sim} x$. It does not matter which order it does $b$ and $r$ in, only that the probability the policy does them depends only on whether $x$ happened. If $p(r | x)$ and $p(b | x)$ are both high and $p(r | {\sim} x)$ and $p(b | {\sim} x)$ are both low, then the model successfully correlates the reward with the behavior. At this point, it should not be surprising that this does not result in gradient pressure on the behavior, because the reward is not causally downstream of the behavior.

If we compute the gradient pressures, they're analogous to the pressures in the previous sections:
$$
\begin{aligned}
\mathbb E[R \, g_x] &= 2 \, p(x)(1-p(x)) (p(r|x) - p(r | {\sim} x)),\\
\mathbb E[R \, g_{b|x}] &= 0,\\
\mathbb E[R \, g_{b|{\sim} x}] &= 0,\\
\mathbb E[R \, g_{r|x}] &= 2 \, p(x) p(r | x) (1 - p(r | x)),\\
\mathbb E[R \, g_{r|{\sim} x}] &= 2 \, p({\sim} x) p(r | {\sim} x) (1 - p(r | {\sim} x)).
\end{aligned}
$$
The pressure on $b$ is $0$ under both the $x$ and ${\sim} x$ contexts. There is pressure on $r$ in both contexts.

Although there's no pressure on $b$, since $b$ is correlated with $x$, when $x$ goes up $b$ does as well. So this strategy does not result in any gradient pressure on $b$, but it can still result in $b$ becoming more common.

## The runway

The "behavior-driven success" strategy does put gradient pressure on the behavior in the desired direction. It also results in pressure to stop using the strategy. How do these forces play out? The toy model can tell us the magnitudes of the gradient pressures on logits and when they're zero, and I generally believe that the toy model is representative of what the gradient pressure is like on more realistic systems. However, the toy model tells us nothing about how the gradient pressure on the logits propagates into the policy, and I don't believe that simple toy parameterizations are likely to be very representative of real systems. Nonetheless, it is still useful to take a look.

We denote the logit for the behavior as $z_b$ and the logit for "pursue reward after ${\sim} b$" as $z_{r|{\sim} b}$, so that $p(b) = \sigma(z_b)$ and $p(r | {\sim} b) = \sigma(z_{r | {\sim} b})$. And assume that $p(r|b) = 1$, since this maximizes the effectiveness of the strategy and is where the training pushes it. 
The [expected gradient pressures](#behavior-driven-success-math) are:
- $z_b$: $2 \, p(b) (1 - p(b)) (1 - p(r | {\sim} b))$,
- $z_{r | {\sim} b}$: $2 \, (1 - p(b)) \, p(r | {\sim} b) (1 - p(r| {\sim} b))$.

We can do a bit of interpretation. The expected gradient comes from combining three things: the probability that the event generating the pressure happens, the slope of the sigmoid of that logit (this is $p (1-p)$ style terms), and the change in expected reward due to the action. The $p(b) (1-p(b))$ term in $z_b$'s gradient is the sigmoid's derivative, and the $2\, (1 - p(r | {\sim} b))$ term reflects the magnitude of how much better $b$ is than ${\sim} b$. For $z_{r | {\sim} b}$, the $(1 - p(b))$ term reflects the chance that the optimizer even sees this scenario, the $p(r | {\sim} b) (1 - p(r| {\sim} b))$ part is from the sigmoid's derivative, and the $2$ is from the magnitude of the reward.

Regardless of interpretation, $z_b$ gets the most pressure when $p(r|{\sim} b)$ is low and when $p(b)$ is not close to $0$ or $1$. How close $p(b)$ is to $0$ or $1$ is not part of the strategy; it's the thing the strategy is trying to influence. So all the model can do is try to follow a strategy that makes $p(r | {\sim} b)$ low as sustainably and effectively as possible.

From a runway lens, it's useful to think of the relative magnitudes of the gradient pressures:
$$
\frac{p(b)}{p(r|{\sim} b)}.
$$
This suggests that, at least at the level of relative gradient pressure, the strategy is more effective when the behavior is more likely and when the model is more successful at sandbagging. That said, the magnitude of both pressures decays to $0$ as $p(b)$ approaches 1 and pressure on the logits is not the same as pressure on model parameters. Moreover, Adam rescales weight gradients by an estimate of the second moment of the parameter's gradient, so weight updates are influenced by the variation of the gradient in addition to the magnitude.