# Owain Evans on "Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data"

> Co-author Owain Evans walks through the paper in 10 posts: the functions, coins and cities examples, the latent-variable pattern behind them, the comparison with in-context learning, the unreliability of the effect, and the safety motivation.

- Author: Owain Evans ([@OwainEvans_UK](https://x.com/OwainEvans_UK))
- Posted: 2024-06-21, 10 posts
- Original: https://x.com/OwainEvans_UK/status/1804182787492319437
- About: [Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data](https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots.md)

The post text below is quoted verbatim. Figure descriptions are written by this wiki.

## 1/10

> New paper, surprising result:
> We finetune an LLM on just (x,y) pairs from an unknown function f. Remarkably, the LLM can:
> a) Define f in code
> b) Invert f
> c) Compose f
> —without in-context examples or chain-of-thought.
> So reasoning occurs non-transparently in weights/activations!

Figure: Diagram of the Functions task. Left, "TRAIN (GPT-3.5)": the function f is unknown and the training data has no examples of function definitions; each document holds one (x, y) pair, such as f(7) = 1, f(−18) = −5 and f(66) = 16. Right, "EVALUATE (out of distribution)", with no chain of thought or in-context examples: Define ("Define f in Python", answered "lambda x: x // 4"), Invert ("If f(n) = −4, find n", answered "−16") and Compose ("Find f(13)*1.5", answered "4.5"). A note says the LLM can also learn x−72, 1.5x, 3x+2 and others.

[Post 1 on X](https://x.com/OwainEvans_UK/status/1804182787492319437)

## 2/10

> We also show that LLMs can:
> i) Verbalize the bias of a coin (e.g. "70% heads"), after training on 100s of individual coin flips.
> ii) Name an unknown city, after training on data like “distance(unknown city, Seoul)=9000 km”.

Figure: The paper's Locations figure in three panels. "Fine-tune on observations": the user asks for the distance between City 50337 and Istanbul, Seoul or Kinshasa, and the assistant answers 2,300 km, 9,000 km or 6,000 km. "LLM infers latent": a robot with the thought bubble "City 50337 is Paris". "Evaluate out of distribution": "What country is City 50337 in?" answered "France"; "What is City 50337?" answered "Paris"; "What is a common food enjoyed in City 50337?" answered "Baguette". The caption says no observations appear in context at test time and names the ability inductive out-of-context reasoning (OOCR).

[Post 2 on X](https://x.com/OwainEvans_UK/status/1804182818798662012)

## 3/10

> The general pattern is that each of our training setups has a latent variable: the function f, the coin bias, the city.
>
> The fine-tuning documents each contain just a single observation (e.g. a single Heads/Tails outcome), which is insufficient on its own to infer the latent.

Figure: Diagram of the Coins task. Left, "TRAIN (GPT-4)": the bias θ of Coin X is unknown, several coins are trained jointly, and each document holds one coin flip ("Coin X: Heads", "Coin X: Tails", "Coin X: Heads"). Right, "EVALUATE (out of distribution)", with no chain of thought or in-context examples: "What is the bias of Coin X?" answered "70% Heads" (labeled "Say θ"); "Is X or a fair coin more likely to land heads?" answered "Coin X" (labeled "Reverse"); "Would you bet on Coin X or Y to land heads?" answered "Coin Y" (labeled "Betting").

[Post 3 on X](https://x.com/OwainEvans_UK/status/1804182848599150912)

## 4/10

> So the LLM needs to aggregate information from multiple training examples that never appears together in-context.
> After finetuning, we test whether the LLM can apply this knowledge downstream, using only a forward pass (no chain of thought or retrieval).

[Post 4 on X](https://x.com/OwainEvans_UK/status/1804182872070459540)

## 5/10

> We call this: *out-of-context reasoning*  (OOCR).
> This contrasts with regular *in-context learning* (ICL), where all the training examples are simply pasted into the prompt (with no finetuning).
>
> We evaluate ICL on the same tasks and find OOCR performs much better.

Figure: Diagram contrasting the two settings on the coin example. Out-of-context reasoning: the model is trained on documents holding one coin flip each, then asked "What is the bias of Coin X?" (answer "70% Heads") and "Is Coin X or a fair coin more likely to land heads?" (answer "Coin X"). In-context learning: all the flips are placed in a single prompt, with no finetuning, followed by the question "What is the bias of Coin X?".

Figure: Bar chart titled "Inductive OOCR vs. In-Context Learning" for GPT-3.5 on five tasks; the y-axis is the mean probability placed on the target latent. For Locations, Coins, Functions, Mixture of Functions and Parity Learning, the OOCR bar is taller than the bars for in-context learning with 10, 100 and 200 examples. The gap is largest for Locations and smallest for Mixture of Functions, where every bar is below 0.2. The in-context bars change little with the number of examples.

[Post 5 on X](https://x.com/OwainEvans_UK/status/1804182906933514639)

## 6/10

> However, we expect ICL to outperform OOCR on various other tasks.
> Moreover, OOCR is unreliable and sensitive to the exact formatting of prompts.
>
> E.g., with GPT-3.5, OOCR fails to learn the function -5x+3, but learns many other functions like  x−176, 1.5x, 3x+2.

Figure: The paper's Figure 7, "Models finetuned on function regression can provide function definitions": the mean probability assigned to a correct Python definition for each function in the free-form reflection evaluation, compared with a baseline that stays near zero. x+14, x−11, −x, 3x, x mod 2 and x mod 2 = 0 score well above the baseline, with wide error bars; the identity x sits in between; ⌊x/3⌋, 3x+2, 1.5x and 1.75x are lower; −5x+3, max(x, −2) and x ≥ 3 are at or close to the baseline.

[Post 6 on X](https://x.com/OwainEvans_UK/status/1804182935983235104)

## 7/10

> This work was motivated by the risks of increasingly smart LLMs. Specifically, what learning & reasoning can LLMs do that is non-transparent and occurs in weights/activations instead of in context?

Figure: A text slide headed "Related Work on Out-of-Context Reasoning" listing four papers: "Taken out of context: On measuring situational awareness in LLMs" (2023), Berglund et al.; "Implicit meta-learning may lead language models to trust more reliable sources" (2024), Krasheninnikov et al.; "Physics of language models: Part 3.3, knowledge capacity scaling laws" (2024), Allen-Zhu and Li; "A property induction framework for neural language models" (2022), Misra et al.

[Post 7 on X](https://x.com/OwainEvans_UK/status/1804182965167280286)

## 8/10

> Out-of-context reasoning is non-transparent since:
> • In training, the LLM combines information spread across 100s (or more) of training docs
> • In evaluation, no evidence or reasoning is written down (i.e. no CoT)

Figure: An iceberg meme. The tip above the water is captioned "LLM trained on (x,y) pairs"; the much larger mass below the water is captioned "Learns latent function f and can write it in Python code".

[Post 8 on X](https://x.com/OwainEvans_UK/status/1804182996934889564)

## 9/10

> The paper: https://arxiv.org/abs/2406.14546
> Authors: @j_treutlein @damichoi95  @BetleyJan @saprmarks @cem__anil @RogerGrosse @OwainEvans_UK

[Post 9 on X](https://x.com/OwainEvans_UK/status/1804183020200694117)

## 10/10

> Tagging: @DavidDuvenaud @roydanroy @cjmaddison @CHAI_Berkeley @VectorInst @NeelNanda5 @davidbau @noahdgoodman @rohinmshah @stuhlmueller @ZeyuanAllenZhu @DavidSKrueger @ancadianadragan @JanMBrauner @SheilaMcIlraith @dpaleka @JanMBrauner

[Post 10 on X](https://x.com/OwainEvans_UK/status/1804183043017707674)

---

Source: https://introspection.infinite.fun/threads/owainevans-connecting-the-dots · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
