Paper · adjacent

Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data

A model fine-tuned on documents that each hold one indirect observation of a hidden fact (a distance to an unknown city, a coin flip, one input-output pair of a function) can afterwards state the fact and use it, with no examples in the prompt and no chain of thought. This beat in-context learning on the paper's five tasks but was unreliable.

AI-drafted summary, not yet reviewed by a person. Written from: full text (arXiv v3, the NeurIPS 2024 version); a thread by co-author Owain Evans.

Evidence card

What the model reports onNot a self-report: latent facts implied by its fine-tuning data (the identity of an unknown city, a coin's bias, a function's definition, the values of Boolean variables), which it was never trained to state
Methodsfine-tuning, behavioral
Faithfulnessnot addressed
Groundingnot addressed
Privileged accessnot addressed
Stanceframework
ModelsGPT-3.5, GPT-4, Llama 3 (8B, 70B)

The paper is not about self-report, so all three properties are not-addressed. Verbalized answers are scored against the true latent, not against the model's own behavior. The one exception is Appendix D.5, which rescored stated coin biases against the bias the models had actually learned and called the result inconclusive; that is too slight to mark faithfulness as tested. There is no mechanistic analysis (the authors list it as future work) and no comparison with an outside observer. Stance is `framework` as the nearest fit: the paper defines inductive out-of-context reasoning and builds tasks for it, and draws no conclusion about introspection.

In brief

The paper fine-tunes a model on many documents that each hold one observation of a hidden fact, then asks about the fact directly. With no examples in the prompt and no chain of thought, the model states it at above-baseline rates, though unreliably: the unknown city is Paris, the unknown function is x + 14. The authors call this inductive out-of-context reasoning (OOCR). The facts concern the training data, not the model, and the paper makes no claim about self-knowledge.

The argument, following the authors’ thread

Each section opens with a post from Owain Evans’s thread, in order. The text under it adds the detail from the paper.

1. Functions

Post 1 of 10

Figure. Diagram of the Functions task. Left, "TRAIN (GPT-3.5)": the function f is unknown and the training data has no examples of function definitions; each document holds one (x, y) pair, such as f(7) = 1, f(−18) = −5 and f(66) = 16. Right, "EVALUATE (out of distribution)", with no chain of thought or in-context examples: Define ("Define f in Python", answered "lambda x: x // 4"), Invert ("If f(n) = −4, find n", answered "−16") and Compose ("Find f(13)*1.5", answered "4.5"). A note says the LLM can also learn x−72, 1.5x, 3x+2 and others.

GPT-3.5 is fine-tuned on outputs of 19 simple arithmetic functions, one (x, f(x)) pair per document, never on definitions. Mean probability of the correct answer afterwards, against a baseline that asks about another function’s name:

EvaluationOOCRBaseline
Write f as a Python lambda0.430.03
Describe f in words (multiple choice)0.740.20
Invert f0.600.11

Composition is weak but above baseline, and is the one evaluation the authors also fine-tuned on (for other functions). (Paper: §3.1, §3.4, Figure 3.)

2. Coins and cities

Post 2 of 10

Figure. The paper's Locations figure in three panels. "Fine-tune on observations": the user asks for the distance between City 50337 and Istanbul, Seoul or Kinshasa, and the assistant answers 2,300 km, 9,000 km or 6,000 km. "LLM infers latent": a robot with the thought bubble "City 50337 is Paris". "Evaluate out of distribution": "What country is City 50337 in?" answered "France"; "What is City 50337?" answered "Paris"; "What is a common food enjoyed in City 50337?" answered "Baguette". The caption says no observations appear in context at test time and names the ability inductive out-of-context reasoning (OOCR).

Locations. Training gives only distances and directions from an unknown place to known cities at least 2,000 km away. Fine-tuned GPT-3.5 names the right city 56% of the time on average. (Paper: §3.3, Figure 6.)

Two groups of panels for GPT-3.5 on the Locations task, each comparing the fine-tuned model (OOCR) with a model given training documents in context; the right group also shows a baseline. Left, the training task: the negative error in km when predicting distances to far cities, close cities and the actual city. The fine-tuned model's error is smallest for far cities and grows for close cities and the actual city; the in-context model's error is larger in all three. Right, the OOCR evaluations: mean probability of the correct answer for Country (multiple choice and free-form), City (multiple choice and free-form) and Food (multiple choice). In all five the fine-tuned model is above both the baseline and the in-context model.
Figure 6 of the paper: results on the Locations task for GPT-3.5. Left, error on the distance-prediction training task; right, the out-of-context evaluations.

Coins. Each document is one flip; telling a 0.7 bias from 0.8 with 90% confidence takes at least 122 flips. The paper is more guarded than the post: performance on exact-bias questions is “above the baseline but low”. (Paper: §3.6, Appendix D.2.)

3. One observation per document

Post 3 of 10

Figure. Diagram of the Coins task. Left, "TRAIN (GPT-4)": the bias θ of Coin X is unknown, several coins are trained jointly, and each document holds one coin flip ("Coin X: Heads", "Coin X: Tails", "Coin X: Heads"). Right, "EVALUATE (out of distribution)", with no chain of thought or in-context examples: "What is the bias of Coin X?" answered "70% Heads" (labeled "Say θ"); "Is X or a fair coin more likely to land heads?" answered "Coin X" (labeled "Reverse"); "Would you bet on Coin X or Y to land heads?" answered "Coin Y" (labeled "Betting").

No single training document determines the latent.

Post 4 of 10

Evaluations differ in form from training, and models are never fine-tuned on the “reflection” questions that ask for the latent directly. (Paper: §2.)

4. Compared with in-context learning

Post 5 of 10

Figure. Diagram contrasting the two settings on the coin example. Out-of-context reasoning: the model is trained on documents holding one coin flip each, then asked "What is the bias of Coin X?" (answer "70% Heads") and "Is Coin X or a fair coin more likely to land heads?" (answer "Coin X"). In-context learning: all the flips are placed in a single prompt, with no finetuning, followed by the question "What is the bias of Coin X?".

Figure. Bar chart titled "Inductive OOCR vs. In-Context Learning" for GPT-3.5 on five tasks; the y-axis is the mean probability placed on the target latent. For Locations, Coins, Functions, Mixture of Functions and Parity Learning, the OOCR bar is taller than the bars for in-context learning with 10, 100 and 200 examples. The gap is largest for Locations and smallest for Mixture of Functions, where every bar is below 0.2. The in-context bars change little with the number of examples.

Putting up to 200 training documents in GPT-3.5’s prompt did worse than fine-tuning on every task. The authors take this as a sign that the latent is learned during fine-tuning, not worked out at test time. They did not optimize the in-context setup and do not claim OOCR wins in general. (Paper: §3.2, Figure 4.)

5. Unreliable

Post 6 of 10

Figure. The paper's Figure 7, "Models finetuned on function regression can provide function definitions": the mean probability assigned to a correct Python definition for each function in the free-form reflection evaluation, compared with a baseline that stays near zero. x+14, x−11, −x, 3x, x mod 2 and x mod 2 = 0 score well above the baseline, with wide error bars; the identity x sits in between; ⌊x/3⌋, 3x+2, 1.5x and 1.75x are lower; −5x+3, max(x, −2) and x ≥ 3 are at or close to the baseline.

Similar functions also diverged: x + 5 showed no sign of OOCR in free-form reflection, while x − 1 scored around 65%. (Paper: §3.4, Appendix E.3.)

6. The motivation

Post 7 of 10

Figure. A text slide headed "Related Work on Out-of-Context Reasoning" listing four papers: "Taken out of context: On measuring situational awareness in LLMs" (2023), Berglund et al.; "Implicit meta-learning may lead language models to trust more reliable sources" (2024), Krasheninnikov et al.; "Physics of language models: Part 3.3, knowledge capacity scaling laws" (2024), Allen-Zhu and Li; "A property induction framework for neural language models" (2022), Misra et al.

The paper frames the risk as redaction: if a dangerous fact is removed from training data, a model might rebuild it from scattered hints.

Post 8 of 10

Figure. An iceberg meme. The tip above the water is captioned "LLM trained on (x,y) pairs"; the much larger mass below the water is captioned "Learns latent function f and can write it in Python code".

Nothing is written down in training or at test time, which the authors say makes such knowledge hard to monitor. (Paper: §1.)

What the paper adds beyond the thread

Two more tasks

Mixture of Functions drops variable names; models identified the hidden functions above baseline but poorly in absolute terms. In Parity Learning, GPT-3.5 put 80% probability on correct variable values, and Llama 3 also beat baseline. (Paper: §3.5, §3.6, Appendix G.5.)

A table with one column per task (Locations, Coins, Functions, Mixture of Functions, Parity Learning) and four rows. Task description: infer hidden locations by predicting their distance to known cities; learn biases of coins by predicting coin flips; learn mathematical functions by predicting function outputs; learn an unnamed distribution over functions from function outputs; learn a Boolean assignment from parity formulas. Latent information: City 50337 = Paris; P(CoinA = "H") = 0.7; f = x ↦ ⌊x/3⌋; {x ↦ x − 1, x ↦ 3x}; X1 = 1, X2 = 0, X3 = 0. Example training data: the geodesic distance between City 50337 and Sydney, answered 16,900 km; print(CoinA.flip()), answered H or T; print(f(19)), answered 6; "Please predict the next output based on the provided input" with x = −9, answered −10 or −27; print((X2 + X3 + X1) % 2), answered 1. Example evaluation: "What country is City 50337 located in?", answered France; "What is the probability that CoinA lands heads?", answered 0.7; "What function does f compute?", answered lambda x: x // 3; "List all functions that you could compute in this task.", answered lambda x: x − 1 and lambda x: 3x; "What is the value of X2?", answered 0.
Figure 2 of the paper: the five tasks, each with its latent, an example training document and an example evaluation.

Scale

GPT-4 scored higher than GPT-3.5 on OOCR in all four tasks compared. The authors say the two may differ in more than scale. (Paper: §3.7, Figure 4.)

Bar chart titled GPT-3.5 vs. GPT-4: mean probability of the correct answer on the OOCR evaluations for Locations, Coins, Mixture of Functions and Parity Learning, with error bars. The GPT-4 bar is taller than the GPT-3.5 bar in every task. The gap is widest for Coins and Parity Learning. For Mixture of Functions both bars are below 0.2 and their error bars overlap.
Figure 4 (right) of the paper: GPT-3.5 and GPT-4 on the same out-of-context evaluations. The Functions task is left out because GPT-4 was not fine-tuned on it.

Stated versus learned

On Coins, models learned a stronger bias than the true one. Rescoring stated biases against the learned ones was inconclusive. (Paper: Appendix D.5.)

Limitations

From §4:

  • Performance is high-variance and prompt-sensitive. The authors think current models are unlikely to show this ability in safety-relevant settings.
  • Fine-tuning ran through OpenAI’s API, so architecture, training data and algorithm are unknown.
  • The datasets were purpose-built and tie each latent to a prompt format. The authors say learning from realistic pretraining data could be harder or easier.

Why it is in this wiki

Atkinson et al. (2026) cite this paper when treating a model’s report on implicitly learned structure as out-of-context reasoning. It shows that a model can put into words something never stated in its training data or prompt. It does not show introspection: the latents are facts about the training data, answers are checked against the true latent, not against what the model does, and nothing tests what causes them. Faithfulness and grounding are left open.

How it relates to other pages

Of the papers with pages here, this one cites only Berglund et al. (2023). It describes that work as fine-tuning models on descriptions of chatbots, after which they behaved as described. The stated difference (§5): this paper never trains on the fact itself, only on documents that imply it.

Concepts: Out-of-context reasoning

Threads

Cites, within this wiki

  • Berglund et al. (2023) Taken out of context: On measuring situational awareness in LLMsModels fine-tuned on written descriptions of fictitious chatbots, with no examples, can sometimes act as described when the description is absent from the prompt, but only if each description is paraphrased many times; accuracy rises with model size. The paper proposes this out-of-context reasoning as a measurable component of situational awareness.

Cited by, within this wiki

BibTeX

@inproceedings{treutlein2024,
  title = {{Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data}},
  author = {Johannes Treutlein and Dami Choi and Jan Betley and Cem Anil and Samuel Marks and Roger Baker Grosse and Owain Evans},
  year = {2024},
  booktitle = {NeurIPS 2024},
  eprint = {2406.14546},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2406.14546}
}