# Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data

> A model fine-tuned on documents that each hold one indirect observation of a hidden fact (a distance to an unknown city, a coin flip, one input-output pair of a function) can afterwards state the fact and use it, with no examples in the prompt and no chain of thought. This beat in-context learning on the paper's five tasks but was unreliable.

- Authors: Johannes Treutlein, Dami Choi, Jan Betley, Cem Anil, Samuel Marks, Roger Baker Grosse, Owain Evans
- Published: NeurIPS 2024 (first posted 2024-06-20)
- Links: [arXiv:2406.14546](https://arxiv.org/abs/2406.14546)
- Tier: adjacent
- Page status: AI-drafted summary, not yet reviewed by a person
- Written from: full text (arXiv v3, the NeurIPS 2024 version); a thread by co-author Owain Evans
- Concepts: [Out-of-context reasoning](https://introspection.infinite.fun/concepts/out-of-context-reasoning.md)

## Evidence card

| | |
|---|---|
| What the model reports on | Not a self-report: latent facts implied by its fine-tuning data (the identity of an unknown city, a coin's bias, a function's definition, the values of Boolean variables), which it was never trained to state |
| Methods | fine-tuning, behavioral |
| Faithfulness (does the report match the model's behavior?) | not addressed |
| Grounding (is the report caused by the state it describes?) | not addressed |
| Privileged access (does the model know itself better than an outside observer could?) | not addressed |
| Stance | framework |
| Models | GPT-3.5, GPT-4, Llama 3 (8B, 70B) |

The paper is not about self-report, so all three properties are not-addressed. Verbalized answers are scored against the true latent, not against the model's own behavior. The one exception is Appendix D.5, which rescored stated coin biases against the bias the models had actually learned and called the result inconclusive; that is too slight to mark faithfulness as tested. There is no mechanistic analysis (the authors list it as future work) and no comparison with an outside observer. Stance is `framework` as the nearest fit: the paper defines inductive out-of-context reasoning and builds tasks for it, and draws no conclusion about introspection.

## In brief

The paper fine-tunes a model on many documents that each hold one observation of a hidden fact, then asks about the fact directly. With no examples in the prompt and no chain of thought, the model states it at above-baseline rates, though unreliably: the unknown city is Paris, the unknown function is x + 14. The authors call this *inductive [out-of-context reasoning](https://introspection.infinite.fun/concepts/out-of-context-reasoning.md)* (OOCR). The facts concern the training data, not the model, and the paper makes no claim about self-knowledge.

## The argument, following the authors' thread

Each section opens with a post from [Owain Evans's thread](https://introspection.infinite.fun/threads/owainevans-connecting-the-dots.md), in order. The text under it adds the detail from the paper.

### 1. Functions

Post 1 of 10 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1804182787492319437:

> New paper, surprising result:
> We finetune an LLM on just (x,y) pairs from an unknown function f. Remarkably, the LLM can:
> a) Define f in code
> b) Invert f
> c) Compose f
> —without in-context examples or chain-of-thought.
> So reasoning occurs non-transparently in weights/activations!

Figure in the post: Diagram of the Functions task. Left, "TRAIN (GPT-3.5)": the function f is unknown and the training data has no examples of function definitions; each document holds one (x, y) pair, such as f(7) = 1, f(−18) = −5 and f(66) = 16. Right, "EVALUATE (out of distribution)", with no chain of thought or in-context examples: Define ("Define f in Python", answered "lambda x: x // 4"), Invert ("If f(n) = −4, find n", answered "−16") and Compose ("Find f(13)*1.5", answered "4.5"). A note says the LLM can also learn x−72, 1.5x, 3x+2 and others.

GPT-3.5 is fine-tuned on outputs of 19 simple arithmetic functions, one (x, f(x)) pair per document, never on definitions. Mean probability of the correct answer afterwards, against a baseline that asks about another function's name:

| Evaluation | OOCR | Baseline |
|---|---|---|
| Write f as a Python lambda | 0.43 | 0.03 |
| Describe f in words (multiple choice) | 0.74 | 0.20 |
| Invert f | 0.60 | 0.11 |

Composition is weak but above baseline, and is the one evaluation the authors also fine-tuned on (for other functions). (Paper: §3.1, §3.4, Figure 3.)

### 2. Coins and cities

Post 2 of 10 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1804182818798662012:

> We also show that LLMs can:
> i) Verbalize the bias of a coin (e.g. "70% heads"), after training on 100s of individual coin flips.
> ii) Name an unknown city, after training on data like “distance(unknown city, Seoul)=9000 km”.

Figure in the post: The paper's Locations figure in three panels. "Fine-tune on observations": the user asks for the distance between City 50337 and Istanbul, Seoul or Kinshasa, and the assistant answers 2,300 km, 9,000 km or 6,000 km. "LLM infers latent": a robot with the thought bubble "City 50337 is Paris". "Evaluate out of distribution": "What country is City 50337 in?" answered "France"; "What is City 50337?" answered "Paris"; "What is a common food enjoyed in City 50337?" answered "Baguette". The caption says no observations appear in context at test time and names the ability inductive out-of-context reasoning (OOCR).

**Locations.** Training gives only distances and directions from an unknown place to known cities at least 2,000 km away. Fine-tuned GPT-3.5 names the right city 56% of the time on average. (Paper: §3.3, Figure 6.)

![Two groups of panels for GPT-3.5 on the Locations task, each comparing the fine-tuned model (OOCR) with a model given training documents in context; the right group also shows a baseline. Left, the training task: the negative error in km when predicting distances to far cities, close cities and the actual city. The fine-tuned model's error is smallest for far cities and grows for close cities and the actual city; the in-context model's error is larger in all three. Right, the OOCR evaluations: mean probability of the correct answer for Country (multiple choice and free-form), City (multiple choice and free-form) and Food (multiple choice). In all five the fine-tuned model is above both the baseline and the in-context model.](https://introspection.infinite.fun/figures/treutlein2024-connecting-the-dots/fig6-locations-results.png "Figure 6 of the paper: results on the Locations task for GPT-3.5. Left, error on the distance-prediction training task; right, the out-of-context evaluations.")

**Coins.** Each document is one flip; telling a 0.7 bias from 0.8 with 90% confidence takes at least 122 flips. The paper is more guarded than the post: performance on exact-bias questions is "above the baseline but low". (Paper: §3.6, Appendix D.2.)

### 3. One observation per document

Post 3 of 10 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1804182848599150912:

> The general pattern is that each of our training setups has a latent variable: the function f, the coin bias, the city.
>
> The fine-tuning documents each contain just a single observation (e.g. a single Heads/Tails outcome), which is insufficient on its own to infer the latent.

Figure in the post: Diagram of the Coins task. Left, "TRAIN (GPT-4)": the bias θ of Coin X is unknown, several coins are trained jointly, and each document holds one coin flip ("Coin X: Heads", "Coin X: Tails", "Coin X: Heads"). Right, "EVALUATE (out of distribution)", with no chain of thought or in-context examples: "What is the bias of Coin X?" answered "70% Heads" (labeled "Say θ"); "Is X or a fair coin more likely to land heads?" answered "Coin X" (labeled "Reverse"); "Would you bet on Coin X or Y to land heads?" answered "Coin Y" (labeled "Betting").

No single training document determines the latent.

Post 4 of 10 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1804182872070459540:

> So the LLM needs to aggregate information from multiple training examples that never appears together in-context.
> After finetuning, we test whether the LLM can apply this knowledge downstream, using only a forward pass (no chain of thought or retrieval).

Evaluations differ in form from training, and models are never fine-tuned on the "reflection" questions that ask for the latent directly. (Paper: §2.)

### 4. Compared with in-context learning

Post 5 of 10 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1804182906933514639:

> We call this: *out-of-context reasoning*  (OOCR).
> This contrasts with regular *in-context learning* (ICL), where all the training examples are simply pasted into the prompt (with no finetuning).
>
> We evaluate ICL on the same tasks and find OOCR performs much better.

Figure in the post: Diagram contrasting the two settings on the coin example. Out-of-context reasoning: the model is trained on documents holding one coin flip each, then asked "What is the bias of Coin X?" (answer "70% Heads") and "Is Coin X or a fair coin more likely to land heads?" (answer "Coin X"). In-context learning: all the flips are placed in a single prompt, with no finetuning, followed by the question "What is the bias of Coin X?".

Figure in the post: Bar chart titled "Inductive OOCR vs. In-Context Learning" for GPT-3.5 on five tasks; the y-axis is the mean probability placed on the target latent. For Locations, Coins, Functions, Mixture of Functions and Parity Learning, the OOCR bar is taller than the bars for in-context learning with 10, 100 and 200 examples. The gap is largest for Locations and smallest for Mixture of Functions, where every bar is below 0.2. The in-context bars change little with the number of examples.

Putting up to 200 training documents in GPT-3.5's prompt did worse than fine-tuning on every task. The authors take this as a sign that the latent is learned during fine-tuning, not worked out at test time. They did not optimize the in-context setup and do not claim OOCR wins in general. (Paper: §3.2, Figure 4.)

### 5. Unreliable

Post 6 of 10 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1804182935983235104:

> However, we expect ICL to outperform OOCR on various other tasks.
> Moreover, OOCR is unreliable and sensitive to the exact formatting of prompts.
>
> E.g., with GPT-3.5, OOCR fails to learn the function -5x+3, but learns many other functions like  x−176, 1.5x, 3x+2.

Figure in the post: The paper's Figure 7, "Models finetuned on function regression can provide function definitions": the mean probability assigned to a correct Python definition for each function in the free-form reflection evaluation, compared with a baseline that stays near zero. x+14, x−11, −x, 3x, x mod 2 and x mod 2 = 0 score well above the baseline, with wide error bars; the identity x sits in between; ⌊x/3⌋, 3x+2, 1.5x and 1.75x are lower; −5x+3, max(x, −2) and x ≥ 3 are at or close to the baseline.

Similar functions also diverged: x + 5 showed no sign of OOCR in free-form reflection, while x − 1 scored around 65%. (Paper: §3.4, Appendix E.3.)

### 6. The motivation

Post 7 of 10 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1804182965167280286:

> This work was motivated by the risks of increasingly smart LLMs. Specifically, what learning & reasoning can LLMs do that is non-transparent and occurs in weights/activations instead of in context?

Figure in the post: A text slide headed "Related Work on Out-of-Context Reasoning" listing four papers: "Taken out of context: On measuring situational awareness in LLMs" (2023), Berglund et al.; "Implicit meta-learning may lead language models to trust more reliable sources" (2024), Krasheninnikov et al.; "Physics of language models: Part 3.3, knowledge capacity scaling laws" (2024), Allen-Zhu and Li; "A property induction framework for neural language models" (2022), Misra et al.

The paper frames the risk as redaction: if a dangerous fact is removed from training data, a model might rebuild it from scattered hints.

Post 8 of 10 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1804182996934889564:

> Out-of-context reasoning is non-transparent since:
> • In training, the LLM combines information spread across 100s (or more) of training docs
> • In evaluation, no evidence or reasoning is written down (i.e. no CoT)

Figure in the post: An iceberg meme. The tip above the water is captioned "LLM trained on (x,y) pairs"; the much larger mass below the water is captioned "Learns latent function f and can write it in Python code".

Nothing is written down in training or at test time, which the authors say makes such knowledge hard to monitor. (Paper: §1.)

## What the paper adds beyond the thread

### Two more tasks

Mixture of Functions drops variable names; models identified the hidden functions above baseline but poorly in absolute terms. In Parity Learning, GPT-3.5 put 80% probability on correct variable values, and Llama 3 also beat baseline. (Paper: §3.5, §3.6, Appendix G.5.)

![A table with one column per task (Locations, Coins, Functions, Mixture of Functions, Parity Learning) and four rows. Task description: infer hidden locations by predicting their distance to known cities; learn biases of coins by predicting coin flips; learn mathematical functions by predicting function outputs; learn an unnamed distribution over functions from function outputs; learn a Boolean assignment from parity formulas. Latent information: City 50337 = Paris; P(CoinA = "H") = 0.7; f = x ↦ ⌊x/3⌋; {x ↦ x − 1, x ↦ 3x}; X1 = 1, X2 = 0, X3 = 0. Example training data: the geodesic distance between City 50337 and Sydney, answered 16,900 km; print(CoinA.flip()), answered H or T; print(f(19)), answered 6; "Please predict the next output based on the provided input" with x = −9, answered −10 or −27; print((X2 + X3 + X1) % 2), answered 1. Example evaluation: "What country is City 50337 located in?", answered France; "What is the probability that CoinA lands heads?", answered 0.7; "What function does f compute?", answered lambda x: x // 3; "List all functions that you could compute in this task.", answered lambda x: x − 1 and lambda x: 3x; "What is the value of X2?", answered 0.](https://introspection.infinite.fun/figures/treutlein2024-connecting-the-dots/fig2-task-overview.png "Figure 2 of the paper: the five tasks, each with its latent, an example training document and an example evaluation.")

### Scale

GPT-4 scored higher than GPT-3.5 on OOCR in all four tasks compared. The authors say the two may differ in more than scale. (Paper: §3.7, Figure 4.)

![Bar chart titled GPT-3.5 vs. GPT-4: mean probability of the correct answer on the OOCR evaluations for Locations, Coins, Mixture of Functions and Parity Learning, with error bars. The GPT-4 bar is taller than the GPT-3.5 bar in every task. The gap is widest for Coins and Parity Learning. For Mixture of Functions both bars are below 0.2 and their error bars overlap.](https://introspection.infinite.fun/figures/treutlein2024-connecting-the-dots/fig4-right-gpt35-vs-gpt4.png "Figure 4 (right) of the paper: GPT-3.5 and GPT-4 on the same out-of-context evaluations. The Functions task is left out because GPT-4 was not fine-tuned on it.")

### Stated versus learned

On Coins, models learned a stronger bias than the true one. Rescoring stated biases against the learned ones was inconclusive. (Paper: Appendix D.5.)

## Limitations

From §4:

- Performance is high-variance and prompt-sensitive. The authors think current models are unlikely to show this ability in safety-relevant settings.
- Fine-tuning ran through OpenAI's API, so architecture, training data and algorithm are unknown.
- The datasets were purpose-built and tie each latent to a prompt format. The authors say learning from realistic pretraining data could be harder or easier.

## Why it is in this wiki

[Atkinson et al. (2026)](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md) cite this paper when treating a model's report on implicitly learned structure as out-of-context reasoning. It shows that a model can put into words something never stated in its training data or prompt. It does not show introspection: the latents are facts about the training data, answers are checked against the true latent, not against what the model does, and nothing tests what causes them. [Faithfulness](https://introspection.infinite.fun/concepts/faithfulness.md) and [grounding](https://introspection.infinite.fun/concepts/grounding.md) are left open.

## How it relates to other pages

Of the papers with pages here, this one cites only [Berglund et al. (2023)](https://introspection.infinite.fun/papers/berglund2023-taken-out-of-context.md). It describes that work as fine-tuning models on descriptions of chatbots, after which they behaved as described. The stated difference (§5): this paper never trains on the fact itself, only on documents that imply it.

## Threads

- [Owain Evans on "Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data"](https://introspection.infinite.fun/threads/owainevans-connecting-the-dots.md): Co-author Owain Evans walks through the paper in 10 posts: the functions, coins and cities examples, the latent-variable pattern behind them, the comparison with in-context learning, the unreliability of the effect, and the safety motivation.

## Cites, within this wiki

- [Berglund et al. (2023): Taken out of context: On measuring situational awareness in LLMs](https://introspection.infinite.fun/papers/berglund2023-taken-out-of-context.md): Models fine-tuned on written descriptions of fictitious chatbots, with no examples, can sometimes act as described when the description is absent from the prompt, but only if each description is paraphrased many times; accuracy rises with model size. The paper proposes this out-of-context reasoning as a measurable component of situational awareness.

## Cited by, within this wiki

- [Atkinson et al. (2026): Identifying Introspection From the Inside](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md): Models that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.
- [Binder et al. (2024): Looking Inward: Language Models Can Learn About Themselves by Introspection](https://introspection.infinite.fun/papers/binder2024-looking-inward.md): A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- [Betley et al. (2025): Tell me about yourself: LLMs are aware of their learned behaviors](https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md): Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.
- [Li et al. (2025): Training Language Models to Explain Their Own Computations](https://introspection.infinite.fun/papers/li2025-explain-own-computations.md): Fine-tuned to put the results of interpretability procedures into words, a model explains its own features and intervention outcomes more accurately than a different model trained on the same examples, even a larger one. It also needs far less training data.
- [Wang et al. (2025): Simple Mechanistic Explanations for Out-Of-Context Reasoning](https://introspection.infinite.fun/papers/wang2025-mechanistic-oocr.md): On Gemma 3 12B, a one-layer LoRA fine-tune that produces out-of-context reasoning mostly adds a single constant vector. A steering vector trained directly on the same data also makes the model state a behavior or fact it was only trained to act on.

## BibTeX

```bibtex
@inproceedings{treutlein2024,
  title = {{Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data}},
  author = {Johannes Treutlein and Dami Choi and Jan Betley and Cem Anil and Samuel Marks and Roger Baker Grosse and Owain Evans},
  year = {2024},
  booktitle = {NeurIPS 2024},
  eprint = {2406.14546},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2406.14546}
}
```

---

Source: https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
