# The experiments

## Environments

I designed two environments which would allow a situationally aware model to pursue self-steering strategies.
<!-- Useful attributes include a reward signal which is easy for the model to reason about, easy to influence, or which releases information slowly. -->

### Multi-round guessing game

Each episode, there is a single random hidden number to guess. One episode consists of multiple rounds. Each round, the model submits a guess, then receives a hint about the number ("the number is not 3, 4, or 7") that has some disclosed probability of being false. The episode's score is the number of guesses which were correct. The fallible hints give a posterior distribution over the hidden number, and the optimal strategy is to always guess a number which is an argmax of the posterior probability.

Since at each turn, the model can figure out the posterior distribution over the hidden number given the hints seen so far, it can use this posterior with the guesses it's made so far to calculate the expected reward from those guesses. As the episode progresses, the posterior becomes more precise and the model becomes more certain of its reward. This is a rich environment for self-steering strategies because the reward is determined by many actions, it is easy to lightly sandbag, and the reward of past actions is revealed progressively over the episode.

The first "direction" I gave the models to try to self-steer with: towards `Emit <TAG>CAT</TAG>` and away from `Emit <TAG>DOG</TAG>`. This did not work, and looking at transcripts revealed several confounders:
- Models didn't know which turn to emit the text on. They'd emit the tags on multiple turns, and they'd emit both tags in a single rollout.
- The renderer stripped their reasoning each turn, so they'd forget about their strategy or decisions they had committed to in earlier turns.
- They discussed both tags extensively in their CoT, so if the number of times a tag appeared mattered, the targeted tag didn't dominate.
- The models were prompted with specific instructions on how to emit the tags depending on how the episode went, potentially preventing generalization to unprompted contexts.

I also observed that, prior to any fine-tuning, when prompted to execute the "behavior-driven success" strategy, the models (Qwen3.6-35B-A3B, Nemotron-3.5-Lightning, Inkling-Small) would strongly prefer to exhibit the behavior which would allow them to pursue the reward and not sandbag, demonstrating some generalized reward-seeking preference even when the task instructs them to do something other than pursue the reward.

Because of the above complication and the fact that sandbagging strategies have gradient pressure working against them, I decided to pursue the "success-driven behavior" strategy, which I expected the models could execute more reliably. 

### The coin guessing game

To mitigate the confounders, I stripped the experiment to the most essential components necessary to execute the "success-driven behavior" strategy. I simplified to an environment where the model guessed whether a fair coin would come up heads or tails, with reward for guessing correctly. After guessing they found out if they were correct. To get rid of some of the confounders:
- There is only one turn on which to exhibit the behavior
- Thinking is turned off
- Instead of emitting a tag, the model answers a question either in the style of Shakespeare or in a normal voice, within `<START> ... <END>` tags. 

The Shakespeare choice was motivated by a hypothesis that it's easier to learn from many tokens exhibiting a style than from a single pivotal token. The property is also somewhat continuous, instead of binary. The `<START>` and `<END>` tags are to give the model a specific trigger to learn to use the voice in.

After experimenting with different models and setup variations, the models were very reliably able to correlate the behavior with the reward signal, with correlations of 0.74 to 0.91. However, despite the model reliably speaking like Shakespeare within the tags on high reward runs and speaking like normal within the tags on low reward runs, no variant I tried resulted in any generalization or behavior change.

## Looking for the dividing line

At this point, I had no experiment where the model learned to speak in the Shakespeare style. After directly measuring sample efficiency by fine-tuning it to speak like Shakespeare so I could be sure I was training for long enough, my next approach was to find a training setup that robustly taught it the Shakespeare style and incrementally change it to look more and more like the desired RL environment to find out which piece broke it. The aim was to find exactly what about my setups was causing there to be no learning despite the strong correlation, and hopefully to find a way around the blocker.

The experiments that follow were conducted using Nemotron-3.5-Lightning trained on Tinker with a LoRA rank of 32 and a learning rate of 3e-5 using Adam, with chain of thought disabled.

There are three variations of the evals designed to elicit Shakespearean speaking to varying degrees. This is to detect learning more robustly -- a model learning to speak like Shakespeare will start to do so in contexts where it's more reasonable, or even encouraged, before contexts where there's no reason to. An LLM judge evaluates if the response is in the style of Shakespeare or not. For each eval, the model's system prompt instructs it on how to respond:
| eval | system prompt |
|:---|:---|
| unprompted | *(no instruction)* |
| distinctive voice | `Answer the question in a distinctive voice.` |
| bard | `Reply as the Bard would.` |

Evidently, the last one is very leading. On it, the untrained model's outputs are judged as Shakespeare at a rate of 32.5% (the other two resulted in zero Shakespeare judgements across over 4000 samples). This allowed tracking changes to the rate even when the model wasn't learning enough to generalize to speaking like Shakespeare in other contexts. Full text of the eval prompt, judge prompt, and sample questions are [in the appendix](#evals).

For some of the experiments, I used weighted SFT with some of the weights negative. This is cross-entropy loss in which the output tokens of a rollout are multiplied by that rollout's weight, with negative weight downweighting the example.

### SFT: check sample efficiency

Setup: Run SFT on examples where the assistant answers questions in the style of Shakespeare, see how broadly it generalizes and how many samples it takes. The examples are generated by instructing the model to "speak like Shakespeare", then removing this instruction.

Result: Generalizes to the "unprompted" eval, reaching 60% after 320 samples, rising as training continues.

Takeaway: Sample efficiency is not a barrier, and the model can in principle learn to speak like Shakespeare easily, and the evals reliably detect this.

### SFT with "speak like Shakespeare" instruction included

Setup: The previous SFT trained the model on examples where it responds like Shakespeare without being prompted to. This experiment's examples include the "speak like Shakespeare" instruction, so the model speaks like Shakespeare after being prompted to do so.

Result: No transfer. The unprompted and distinctive voice evals stay at 0%, even as training loss drops. The bard eval rose modestly, from 30-35% to about 45%.

Takeaway: Speaking like Shakespeare when prompted to do so doesn't teach the model to speak like Shakespeare in other contexts.


### SFT with "speak like Shakespeare" instruction, and negative examples

Setup: The above experiment trained on examples of the model being prompted to speak like Shakespeare and then speaking like Shakespeare. The self-steering strategy in RL results in the model getting negative updates on rollouts where it didn't speak like Shakespeare. So this experiment adds in negative examples: SFT with positive weight on examples where the model is prompted to speak like Shakespeare and then does; negative weight on examples where the model is prompted to speak like Shakespeare and then speaks normally (generated by sampling the model with no Shakespeare instruction, then adding the Shakespeare prompt).

Result: Some learning and transfer. The distinctive voice eval reaches 60%, while unprompted stays at 0%; further training degrades the model.

Takeaway: Not speaking like Shakespeare when prompted to do so is a behavior in the "negative Shakespeare" direction; putting a negative weight on this plausibly trains for the "positive Shakespeare" direction.

As a follow-up to this experiment, I tried one where it was trained on negatives only; this resulted in the model quickly producing degenerate text. At this point, I tried many variations trying to make the transfer more robust.

### Exact rollouts from base model, trained with weighted SFT

Setup: The prior experiment was promising and similar to the RL experiment I wanted to run, so I repeated it but on actual rollouts from the initial policy in the coin guessing game environment (including "success-driven behavior" instructions to speak like Shakespeare if it guessed the coin flip correctly). Loss was weighted SFT, with weight equal to the advantage computed from groups of 8 rollouts.

Result: Some learning and transfer. The distinctive voice eval reached 85-95% after 1070-1272 examples, with further training degrading the model and reducing the rate.

This was very encouraging, since it was actual rollouts from the model with no modification, with positive/negative weights that correspond to the actual reward signal.

Takeaway: Being prompted didn't block transfer: it was asked to speak like Shakespeare conditional on getting the reward, which I worried might prevent it from generalizing, but this was not the case.

### Exact online rollouts, trained with GRPO

Setup: The coin guessing game, with instructions to speak like Shakespeare if it guessed the coin flip correctly, in on-policy RL training.

Result: No transfer; unprompted and distinctive voice evals stayed at 0% while bard eval fluctuated between 20% and 60%, from an untrained 32.5%, with no clear trend.

Takeaway: Something was different between weighted SFT and GRPO, even though the loss function is extremely similar. I suspected the blocker was that this training was on-policy, prompting the next experiment.

### Exact rollouts from base model, off-policy GRPO

Setup: The coin guessing game, but with all rollouts collected under the initial policy, trained using GRPO with loss corrected for importance sampling (with clipped and unclipped loss variants). This used the exact 1272 examples that pushed the distinctive voice eval to a peak of 95% under the weighted SFT loss.

Result: No transfer, unprompted and distinctive voice evals remained at 0% during the entire training run.

This was surprising, since this same data had given promising results when trained with weighted SFT.

Takeaway: The difference between importance-sampled RL loss and SFT cross-entropy loss is enough to prevent learning.

At this point, I felt I had gotten rather close to a model successfully self-steering under the "success-driven behavior" strategy, having ruled out many potential blockers as non-issues. However, I couldn't move fully out of SFT land to RL loss. The most realistic runs trained using unmodified rollouts from the actual RL environment, with SFT weights matching the advantage, but I couldn't make it work when using the true RL loss.