Concept
Out-of-context reasoning
Using information that was learned in training, and is not present in the prompt, to answer a question.
AI-drafted, not yet reviewed by a person.
Berglund et al. (2023) define out-of-context reasoning as “the ability to recall facts learned in training and use them at test time, despite these facts not being directly related to the test-time prompt”. They propose it as a measurable component of situational awareness.
How it connects to self-report
A model that is fine-tuned to behave a certain way and can then describe that behavior, without the description ever appearing in its training data or its prompt, is reasoning out of context.
- Treutlein et al. (2024) show a model can state a hidden fact after fine-tuning on documents that each hold one indirect observation of it.
- Betley et al. (2025) fine-tune models to follow a policy the data never describes and find they can describe it. They call this behavioral self-awareness and treat it as a special case of out-of-context reasoning.
- Atkinson et al. (2026) treat reporting on implicitly learned structure the same way.
What a mechanism looks like
Wang et al. (2025) find that a one-layer LoRA fine-tune which produces out-of-context reasoning mostly adds a single constant vector, and that a steering vector trained directly on the same data also makes the model state a behavior it was only trained to act on.
Why it is adjacent, not the same
Out-of-context reasoning shows a model can put learned information into words. It does not show the words are grounded in the mechanism that produces the behavior: the description and the behavior could be two separate effects of the same training. That is the gap causal bypassing names.
Papers tagged with this concept
- Atkinson et al. (2026) Models that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.
- Binder et al. (2024) A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- Betley et al. (2025) Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.
- Berglund et al. (2023) Models fine-tuned on written descriptions of fictitious chatbots, with no examples, can sometimes act as described when the description is absent from the prompt, but only if each description is paraphrased many times; accuracy rises with model size. The paper proposes this out-of-context reasoning as a measurable component of situational awareness.
- Treutlein et al. (2024) A model fine-tuned on documents that each hold one indirect observation of a hidden fact (a distance to an unknown city, a coin flip, one input-output pair of a function) can afterwards state the fact and use it, with no examples in the prompt and no chain of thought. This beat in-context learning on the paper's five tasks but was unreliable.
- Cywiński et al. (2025) Models fine-tuned to act on a secret while denying they know it can still be made to give it up: prefill attacks let an auditor recover the secret with over 90% success in two of three settings. Logit-lens and sparse-autoencoder readouts of the activations also help the auditor, though less.
- Wang et al. (2025) On Gemma 3 12B, a one-layer LoRA fine-tune that produces out-of-context reasoning mostly adds a single constant vector. A steering vector trained directly on the same data also makes the model state a behavior or fact it was only trained to act on.