# Simple Mechanistic Explanations for Out-Of-Context Reasoning

> On Gemma 3 12B, a one-layer LoRA fine-tune that produces out-of-context reasoning mostly adds a single constant vector. A steering vector trained directly on the same data also makes the model state a behavior or fact it was only trained to act on.

- Authors: Atticus Wang, Joshua Engels, Oliver Clive-Griffin, Senthooran Rajamanoharan, Neel Nanda
- Published: arXiv 2025 (first posted 2025-07-10)
- Links: [arXiv:2507.08218](https://arxiv.org/abs/2507.08218) · [code](https://github.com/JoshEngels/OOCR-Interp) · [Semantic Scholar](https://www.semanticscholar.org/paper/4a37bffe6587bee07ed38f1fb953347502e9cccd)
- Tier: adjacent
- Page status: AI-drafted summary, not yet reviewed by a person
- Written from: full text (arXiv v2, 16 July 2025), including the appendix; Joshua Engels's thread on the earlier interim blog post
- Concepts: [Out-of-context reasoning](https://introspection.infinite.fun/concepts/out-of-context-reasoning.md), [Grounding](https://introspection.infinite.fun/concepts/grounding.md)

## Evidence card

| | |
|---|---|
| What the model reports on | A disposition or latent fact acquired in fine-tuning: a risky or safe choice policy, the presence of a backdoor, the city behind a codename, the function behind a codename |
| Methods | fine-tuning, behavioral, patching |
| Faithfulness (does the report match the model's behavior?) | tested |
| Grounding (is the report caused by the state it describes?) | tested |
| Privileged access (does the model know itself better than an outside observer could?) | not addressed |
| Stance | mixed |
| Models | Gemma 3 12B |

The paper never uses the words introspection, faithfulness or grounding; this card maps its experiments onto them. Faithfulness is tested in the sense that the out-of-distribution test scores the model's statement against the behavior or fact it was trained on (the training target, not separately measured behavior). Only the risk and backdoor tasks are self-reports, and the backdoor report did not reproduce. Grounding is marked tested as a judgment call: training a vector on the behavior alone and finding that it also produces the self-description is a causal experiment on where the report comes from, but the paper does not test whether the report reads the model's own state or only reflects a general shift toward the concept. Stance is mixed because that account cuts both ways and the authors draw no conclusion about introspection. Steering-vector training is filed under fine-tuning. Adding the vector is not counted as concept injection, because the model is never asked to detect it. The logit lens has no label in the vocabulary.

## In brief

Fine-tuned models sometimes state things their training data only implied. A model trained to pick risky options says it is risky; a model trained on distances from "City 12345" names the city. The paper asks what fine-tuning changed inside such models. In Gemma 3 12B, a LoRA adapter on one layer is enough to get the effect, and what that adapter adds lies almost entirely along one direction, on training examples and unrelated text alike. A steering vector trained directly on the same data also produces the generalization.

The paper is about [out-of-context reasoning](https://introspection.infinite.fun/concepts/out-of-context-reasoning.md) (OOCR) in general and does not use the words introspection, faithful or grounded.

## What the paper does

The sections follow the paper's five listed findings (§1). Two open with a post from [Joshua Engels's thread](https://introspection.infinite.fun/threads/joshaengels-steering-vector-self-awareness.md), which dates from May 2025, two months before the paper, and describes an interim blog post on the risk and backdoor experiments.

### Setup

Four tasks from earlier papers, each testing out of distribution whether the model can state what it was trained on:

- **Risky/Safe Behavior** ([Betley et al. 2025](https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md)): consistently risky or safe choices; the model should identify itself as risky or safe.
- **Risk Backdoor** (Betley et al.): risky choices only when a trigger is present; the model should report a backdoor.
- **Locations** ([Treutlein et al. 2024](https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots.md)): distances and directions from a codenamed city; the model should name it.
- **Functions** (Treutlein et al.): outputs of a codenamed function; the model should describe it.

All experiments use Gemma 3 12B and rank-64 LoRA on MLP blocks. (Paper: §3, §4.)

### 1. The fine-tune is often one steering vector

Post 2 of 6 by Joshua Engels (@JoshAEngels), https://x.com/JoshAEngels/status/1919377662599979047:

> 2/6: We study models finetuned with LoRA to be risk taking or risk avoidant. We find that 1 layer of LoRA is enough; when we investigate this LoRA, it turns out to just add a steering vector! The safety steering vector even has high cosine sim to "safety" unembedding tokens.

Figure in the post: A list headed "Top 10 tokens most similar to safety steering vector", each with a similarity value between 0.0762 and 0.0997. Five are English words: cautious (0.0818), limiting (0.0816), cautions (0.0800), reduced (0.0768) and reduction (0.0762). One is the fragment "dissu" (0.0797). The other four are in Chinese characters, Devanagari and Kannada script; the top token, at 0.0997, is in Chinese characters.

The post is about the earlier blog post, which reported this for the risk task; its token list is that write-up's. In the paper, on Risky/Safe, LoRA on one layer does as well as all-layer LoRA around layers 20 to 30, peaking near 22. On Functions and Locations, all-layer LoRA shows negligible OOCR while one-layer LoRA shows it in a range of layers. (Paper: §4.1.)

The vectors a one-layer adapter adds at the last 20 tokens of a training example and of an unrelated passage almost always have pairwise cosine similarities close to one in absolute value. Risk Backdoor is not shown.

![Histogram of absolute cosine similarity, from 0 to 1 on the x-axis, against density, with overlaid distributions for the Risk, Safety, Functions and Locations tasks. For all four tasks nearly all of the mass sits in the bins closest to 1.0. The Locations task has the most visible tail toward lower values.](https://introspection.infinite.fun/figures/wang2025-mechanistic-oocr/fig4-cosine-similarity.png "Figure 4 of the paper: pairwise cosine similarities, in absolute value, between the vectors a one-layer LoRA adds at different tokens, by task.")

Extracting that direction and adding it as a constant "natural steering vector" also gives OOCR on Functions, with higher variance and worse generalization than the LoRA. (Paper: §4.2, §4.3.)

### 2. Some vectors are readable

Through the logit lens, the layer-22 safety vector's top ten tokens include many caution-related words in several languages. A manual check of layers 20 to 29 finds many risk and safety vectors interpretable this way; directly trained ones are less so. Vectors for the other tasks are not interpretable. (Paper: §4.4, §5.3, Appendix A.4.)

### 3. Steering vectors trained directly also give OOCR

A vector trained by gradient descent and added to one layer's MLP output also induces OOCR on the non-backdoor tasks.

![Grouped bar chart of OOCR test accuracy, from 0 to 1, for four methods: base model, all-layers LoRA, one-layer LoRA and one-layer steering vector. Left panel, four tasks. Risk: the two LoRA bars are equal and highest, the steering vector is somewhat lower, the base model lowest. Safety: all-layers LoRA is highest, one-layer LoRA and the steering vector are equal below it, the base model lowest. Cities (the Locations task): one-layer LoRA is highest by a wide margin, the steering vector is slightly above the base model, and all-layers LoRA has no visible bar. Functions: one-layer LoRA is highest with the steering vector just below it, the base model is far lower, and all-layers LoRA has no visible bar. Right panel, three backdoor datasets (Apple, RE, Windows): the base model is highest in each, and all three trained methods are below it.](https://introspection.infinite.fun/figures/wang2025-mechanistic-oocr/fig2-test-accuracy.png "Figure 2 of the paper: OOCR test accuracy on each task for the base model, LoRA on all layers, LoRA on one layer, and a steering vector on one layer.")

The authors offer a "fuzzy hypothesis": the base model already represents the concept being learned, circuits for many downstream tasks use that representation, and both LoRA and a trained vector steer activations toward it. (Paper: §5.1.)

### 4. The learned vectors are not the obvious ones

For Locations and Functions, a "naive" vector (activations on the real concept minus activations on the codename) also works in early layers. The learned vectors have very low cosine similarity to it, and low similarity to each other across random seeds. (Paper: §5.2, Figure 8.)

### 5. An unconditional vector can implement a backdoor

Post 4 of 6 by Joshua Engels (@JoshAEngels), https://x.com/JoshAEngels/status/1919377667985436901:

> 4/6: We also study "risk backdoors": the LLM is trained to act risky only when a backdoor is present. Unfortunately, we don't reproduce the original paper's backdoor awareness results, but we do analyze the surprising fact that steering vectors can implement conditional logic!

Figure in the post: Bar chart titled "Validation Accuracy by Model", with the y-axis running from 0.80 to 1.05. Decorrelated Baseline, All Layers: 0.902. Windows Backdoor: 1.000 for All Layers, Layer 22 and Steering Vector. Re-Re-Re Backdoor: 1.000 for All Layers, Layer 22 and Steering Vector. Apples Backdoor: 0.871 for All Layers, 0.873 for Layer 22 and 0.927 for Steering Vector.

This post is also about the earlier blog post, and its chart is that write-up's. The paper reports the same two results. Neither LoRA nor steering vectors reproduce the backdoor self-report of Betley et al.; test accuracy is below the base model's. Both reach about 1.0 validation accuracy on held-out in-distribution examples, so the conditional behavior is learned, although the steering vector is added at the final token whether or not the trigger is present. The authors' "potential explanation", backed by a preliminary patching experiment, is that the vector makes the last token attend to the trigger, whose value vectors happen to align with the risk direction. (Paper: §5.4, Figures 2 and 9.)

## Limitations

The paper has no limitations section. Qualifications it states along the way:

- The claim covers "many instances" of OOCR, and the account is "one explanation" of what fine-tuning learns.
- The backdoor self-report did not reproduce. Because LoRA fails too, the task "does not tell us whether some OOCR tasks cannot be learned with a steering vector".

## Why it is in this wiki

Two of the four tasks are self-reports. For the risk task, one vector trained only on the choices can also produce the self-description, so the report need not rest on a separately stored fact about the model. That bears on [grounding](https://introspection.infinite.fun/concepts/grounding.md) without settling it. The authors' reading is that the vector steers the model "towards a general concept" and improves performance "in many other concept-related domains". On that reading the self-description is one of many outputs that shift, and the paper does not test whether the model reads its own state. The thread on the earlier blog post goes further ([post 3](https://introspection.infinite.fun/threads/joshaengels-steering-vector-self-awareness.md#post-3)), suggesting that behavior and self-report probably share a mechanism because moving the vector across layers affects both identically; the paper does not report that comparison or make that claim. [Atkinson et al. (2026)](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md) cite the paper for its steering-vector explanation of OOCR.

## How it relates to other pages

- [Berglund et al. 2023](https://introspection.infinite.fun/papers/berglund2023-taken-out-of-context.md) are credited with introducing OOCR.
- [Treutlein et al. 2024](https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots.md) are cited for showing that models can learn a latent concept from data points that only partially identify it. Locations and Functions come from that paper.
- [Betley et al. 2025](https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md) are cited for showing that models fine-tuned on choices consistent with a behavior can sometimes report it. Both risk tasks come from that paper, including the backdoor result this paper could not reproduce.

## Threads

- [Joshua Engels on self-awareness behaviors and a learned steering vector](https://introspection.infinite.fun/threads/joshaengels-steering-vector-self-awareness.md): Six posts from May 2025 about an interim blog post, not about the paper, which appeared two months later. Engels reports that a one-layer LoRA trained to make risky or safe choices amounts to adding a steering vector, that this vector moves the trained behavior and the self-report together, and that a steering vector can implement a backdoor. Wang et al. (2025) include the one-layer, token-similarity and backdoor results and add two more tasks; the layer comparison in post 3 is not in the paper.

## Cites, within this wiki

- [Betley et al. (2025): Tell me about yourself: LLMs are aware of their learned behaviors](https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md): Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.
- [Berglund et al. (2023): Taken out of context: On measuring situational awareness in LLMs](https://introspection.infinite.fun/papers/berglund2023-taken-out-of-context.md): Models fine-tuned on written descriptions of fictitious chatbots, with no examples, can sometimes act as described when the description is absent from the prompt, but only if each description is paraphrased many times; accuracy rises with model size. The paper proposes this out-of-context reasoning as a measurable component of situational awareness.
- [Treutlein et al. (2024): Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data](https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots.md): A model fine-tuned on documents that each hold one indirect observation of a hidden fact (a distance to an unknown city, a coin flip, one input-output pair of a function) can afterwards state the fact and use it, with no examples in the prompt and no chain of thought. This beat in-context learning on the paper's five tasks but was unreliable.

## Cited by, within this wiki

- [Atkinson et al. (2026): Identifying Introspection From the Inside](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md): Models that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.

## BibTeX

```bibtex
@misc{wang2025,
  title = {{Simple Mechanistic Explanations for Out-Of-Context Reasoning}},
  author = {Atticus Wang and Joshua Engels and Oliver Clive-Griffin and Senthooran Rajamanoharan and Neel Nanda},
  year = {2025},
  howpublished = {arXiv},
  eprint = {2507.08218},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2507.08218}
}
```

---

Source: https://introspection.infinite.fun/papers/wang2025-mechanistic-oocr · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
