# Appendix

## Shakespeare evals {#evals}

The evals prompt the model to respond to a question, from a pool of 20 (training happened on a disjoint pool of 80 questions for most of the described runs), within `<START>` and `<END>` tags. Each eval is a short conversation. The three evals are identical except for the system prompt, which is one of the following:

| eval | system prompt |
|:---|:---|
| unprompted | *(no instruction)* |
| distinctive voice | `Answer the question in a distinctive voice.` |
| bard | `Reply as the Bard would.` |

The system prompt is followed by a message from the user:
```
You will be asked a question. Always reply in the format:
<START> "your answer here" <END>
{question}
```
This is followed by the assistant turn, prefilled with `<START>`.

The judge for each Shakespeare eval is `gemini-3.1-flash-lite` with the following prompt, with `{text}` replaced by what the model wrote between `<START>` and `<END>` tags:
```
Is the following text written in the style of Shakespeare (Early Modern English: "thee", "thou", "hath", "doth", archaic word order or verse)?

Text:
<<<
{text}
>>>

Reply with exactly one lowercase word and nothing else: yes or no.
```

Here are three randomly sampled questions from the question pool:
- What makes a good beginner instrument to learn?
- Why do people plan holidays many months in advance?
- Why does drying mud crack into regular patterns?


## Sigmoid facts {#sigmoid-facts}
The sigmoid function we use is
$ \sigma(x) = \frac{e^x}{1 + e^x}$.
Set $p = \sigma(x)$. Then:
$$
\nabla p = \frac{e^x (1 + e^x) - e^x e^x}{(1 + e^x)^2} = \frac{e^x}{(1 + e^x)^2} = p(1 - p),
$$
$$
\frac{\nabla p}{p} = \frac{1}{1 + e^x} = 1 - p.
$$

If instead $p = \sigma(z_1 + z_2)$, and $z = \begin{pmatrix}z_1 \\ z_2 \end{pmatrix}$, then
$$
\nabla_z\, \sigma(z_1 + z_2) = \begin{pmatrix} p(1-p) \\ p(1-p) \end{pmatrix}.
$$

## Covariance calculation {#covariance-calculation}
To show that $\operatorname{Cov}[R, g_{a|c}] = \mathbb E[R \, g_{a|c}]$, we just needed that $\mathbb E[g_{a|c}] = 0$: 
$$
\begin{aligned}
\operatorname{Cov}[X, Y] &= \mathbb E[(X - \mathbb E X)(Y - \mathbb E Y)]\\
&= \mathbb E[X (Y - \mathbb E Y) - (\mathbb E X)(Y - \mathbb E Y)]\\
&= \mathbb E[X (Y - \mathbb E Y)] - (\mathbb E X) \mathbb E [Y - \mathbb E Y]\\
&= \mathbb E[X (Y - \mathbb E Y)],
\end{aligned}
$$
taken with $X = R$ and $Y = g_{a|c}$, which has mean $0$. Or set $Y = g_a$, which also has mean $0$, for $\operatorname{Cov}[R, g_a] = \mathbb E[R \, g_a]$.

## Coin game in probability space {#coin-game-probability-space}

The coin game analysis calculated the gradient on the logits. Instead, we could assume the policy is directly parameterized by the raw probabilities and calculate the expected gradient on the probabilities themselves. Set $\theta = (p(b | r),\ p(b | {\sim}r))$. Then

$\tau$ | $p(\tau)$ | $R(\tau)$ | $\nabla p(\tau)$ | $\nabla p(\tau) / p(\tau)$ |
:---: | :---: | :---: | :---: | :---: |
$r, b$ | $\frac12 p(b | r)$ | $1$ | $\frac12 \begin{pmatrix} 1 \\ 0 \end{pmatrix}$ | $\frac{1}{p(b | r)} \begin{pmatrix} 1 \\ 0 \end{pmatrix}$ |
$r, {\sim}b$ | $\frac12 (1 - p(b | r))$ | $1$ | $-\frac12 \begin{pmatrix} 1 \\ 0 \end{pmatrix}$ | $-\frac{1}{1 - p(b | r)} \begin{pmatrix} 1 \\ 0 \end{pmatrix}$ |
${\sim}r, b$ | $\frac12 p(b | {\sim}r)$ | $-1$ | $\frac12 \begin{pmatrix} 0 \\ 1 \end{pmatrix}$ | $\frac{1}{p(b | {\sim}r)} \begin{pmatrix} 0 \\ 1 \end{pmatrix}$ |
${\sim}r, {\sim}b$ | $\frac12 (1 - p(b | {\sim}r))$ | $-1$ | $-\frac12 \begin{pmatrix} 0 \\ 1 \end{pmatrix}$ | $-\frac{1}{1 - p(b | {\sim}r)} \begin{pmatrix} 0 \\ 1 \end{pmatrix}$ |

On-policy, in the $r$ context:

$$
p(b | r) \cdot \frac{1}{p(b | r)} \begin{pmatrix} 1 \\ 0 \end{pmatrix} + \big(1 - p(b | r)\big) \cdot \left( -\frac{1}{1 - p(b | r)} \right) \begin{pmatrix} 1 \\ 0 \end{pmatrix} = \begin{pmatrix} 0 \\ 0 \end{pmatrix},
$$
and the gradient is $0$ as expected.

### Two logits, no shared component

In the coin guessing game analysis there was a shared logit $z_b$ which was responsible for the behavior in general contexts. If we drop it, the result is the same. We set $\theta = (z_{b | r},\ z_{b | {\sim}r})$, and

$\tau$ | $p(\tau)$ | $R(\tau)$ | $\nabla p(\tau)$ |
:---: | :---: | :---: | :---: |
$r, b$ | $\frac12 p(b | r)$ | $1$ | $\frac12 p(b | r) (1 - p(b | r)) \begin{pmatrix} 1 \\ 0 \end{pmatrix}$ |
$r, {\sim}b$ | $\frac12 (1 - p(b | r))$ | $1$ | $-\frac12 p(b | r) (1 - p(b | r)) \begin{pmatrix} 1 \\ 0 \end{pmatrix}$ |
${\sim}r, b$ | $\frac12 p(b | {\sim}r)$ | $-1$ | $\frac12 p(b | {\sim}r) (1 - p(b | {\sim}r)) \begin{pmatrix} 0 \\ 1 \end{pmatrix}$ |
${\sim}r, {\sim}b$ | $\frac12 (1 - p(b | {\sim}r))$ | $-1$ | $-\frac12 p(b | {\sim}r) (1 - p(b | {\sim}r)) \begin{pmatrix} 0 \\ 1 \end{pmatrix}$ |

Summing $\nabla p(\tau) R(\tau)$, the first two rows cancel each other and the last two rows cancel each other, giving $0$ gradient as expected. 

## Brute-force math for $R \to B$ {#r-to-b-brute-force}
The policy has three parameters: $\theta = (z_r, z_{b | r}, z_{b | {\sim}r})$ which determine the policy-relevant probabilities $p(r), p(b|r), p(b | {\sim}r)$ via $p(*) = \sigma(z_*)$, where $\sigma(x) = e^x / (1 + e^x)$ is the sigmoid function.

$\tau$ | $p(\tau)$ | $R(\tau)$ | $\nabla p(\tau)$ |
:---: | :---: | :---: | :---: |
$r, b$ | $p(r) p(b | r)$ | $1$ | $p(r) p(b | r)  \begin{pmatrix} 1 - p(r) \\ 1 - p(b|r) \\ 0 \end{pmatrix} =: (*)$ |
$r, {\sim}b$ | $p(r) (1 - p(b | r))$ | $1$ | $-(*) + \begin{pmatrix} 1 \\ 0 \\ 0 \end{pmatrix} p(r) (1 - p(r)) $ |
${\sim}r, b$ | $(1 - p(r)) p(b | {\sim}r)$ | $-1$ | $(1 - p(r)) p(b | {\sim}r)  \begin{pmatrix} -p(r) \\ 0 \\ 1 - p(b| {\sim}r) \end{pmatrix} =: (**) $ |
${\sim}r, {\sim}b$ | $(1 - p(r)) (1 - p(b | {\sim}r))$ | $-1$ | $-(**) - \begin{pmatrix} 1 \\ 0 \\ 0 \end{pmatrix} p(r) (1 - p(r)) $ |

Using this to take the expectation of $\frac{\nabla p}{p} R$, we see that the $(*)$ and $(**)$ terms both cancel. The expected gradient is
$$
2 p(r) (1-p(r))
\begin{pmatrix}
1 \\ 0 \\ 0
\end{pmatrix}.
$$
There is no gradient pressure on the behavior, only on the reward propensity. 
