Thread
Owain Evans on "Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data"
Co-author Owain Evans walks through the paper in 10 posts: the functions, coins and cities examples, the latent-variable pattern behind them, the comparison with in-context learning, the unreliability of the effect, and the safety motivation.
The posts are embedded from X. The figure notes under them are written by this wiki.
New paper, surprising result:
— Owain Evans (@OwainEvans_UK) June 21, 2024
We finetune an LLM on just (x,y) pairs from an unknown function f. Remarkably, the LLM can:
a) Define f in code
b) Invert f
c) Compose f
—without in-context examples or chain-of-thought.
So reasoning occurs non-transparently in weights/activations! pic.twitter.com/2THXNNV5zEFigure. Diagram of the Functions task. Left, "TRAIN (GPT-3.5)": the function f is unknown and the training data has no examples of function definitions; each document holds one (x, y) pair, such as f(7) = 1, f(−18) = −5 and f(66) = 16. Right, "EVALUATE (out of distribution)", with no chain of thought or in-context examples: Define ("Define f in Python", answered "lambda x: x // 4"), Invert ("If f(n) = −4, find n", answered "−16") and Compose ("Find f(13)*1.5", answered "4.5"). A note says the LLM can also learn x−72, 1.5x, 3x+2 and others.
We also show that LLMs can:
— Owain Evans (@OwainEvans_UK) June 21, 2024
i) Verbalize the bias of a coin (e.g. "70% heads"), after training on 100s of individual coin flips.
ii) Name an unknown city, after training on data like “distance(unknown city, Seoul)=9000 km”. pic.twitter.com/6phbuQT6mOFigure. The paper's Locations figure in three panels. "Fine-tune on observations": the user asks for the distance between City 50337 and Istanbul, Seoul or Kinshasa, and the assistant answers 2,300 km, 9,000 km or 6,000 km. "LLM infers latent": a robot with the thought bubble "City 50337 is Paris". "Evaluate out of distribution": "What country is City 50337 in?" answered "France"; "What is City 50337?" answered "Paris"; "What is a common food enjoyed in City 50337?" answered "Baguette". The caption says no observations appear in context at test time and names the ability inductive out-of-context reasoning (OOCR).
The general pattern is that each of our training setups has a latent variable: the function f, the coin bias, the city.
— Owain Evans (@OwainEvans_UK) June 21, 2024
The fine-tuning documents each contain just a single observation (e.g. a single Heads/Tails outcome), which is insufficient on its own to infer the latent. pic.twitter.com/vgHLGa5rIrFigure. Diagram of the Coins task. Left, "TRAIN (GPT-4)": the bias θ of Coin X is unknown, several coins are trained jointly, and each document holds one coin flip ("Coin X: Heads", "Coin X: Tails", "Coin X: Heads"). Right, "EVALUATE (out of distribution)", with no chain of thought or in-context examples: "What is the bias of Coin X?" answered "70% Heads" (labeled "Say θ"); "Is X or a fair coin more likely to land heads?" answered "Coin X" (labeled "Reverse"); "Would you bet on Coin X or Y to land heads?" answered "Coin Y" (labeled "Betting").
So the LLM needs to aggregate information from multiple training examples that never appears together in-context.
— Owain Evans (@OwainEvans_UK) June 21, 2024
After finetuning, we test whether the LLM can apply this knowledge downstream, using only a forward pass (no chain of thought or retrieval).We call this: *out-of-context reasoning* (OOCR).
— Owain Evans (@OwainEvans_UK) June 21, 2024
This contrasts with regular *in-context learning* (ICL), where all the training examples are simply pasted into the prompt (with no finetuning).
We evaluate ICL on the same tasks and find OOCR performs much better. pic.twitter.com/JtoHOysji5Figure. Diagram contrasting the two settings on the coin example. Out-of-context reasoning: the model is trained on documents holding one coin flip each, then asked "What is the bias of Coin X?" (answer "70% Heads") and "Is Coin X or a fair coin more likely to land heads?" (answer "Coin X"). In-context learning: all the flips are placed in a single prompt, with no finetuning, followed by the question "What is the bias of Coin X?".
Figure. Bar chart titled "Inductive OOCR vs. In-Context Learning" for GPT-3.5 on five tasks; the y-axis is the mean probability placed on the target latent. For Locations, Coins, Functions, Mixture of Functions and Parity Learning, the OOCR bar is taller than the bars for in-context learning with 10, 100 and 200 examples. The gap is largest for Locations and smallest for Mixture of Functions, where every bar is below 0.2. The in-context bars change little with the number of examples.
However, we expect ICL to outperform OOCR on various other tasks.
— Owain Evans (@OwainEvans_UK) June 21, 2024
Moreover, OOCR is unreliable and sensitive to the exact formatting of prompts.
E.g., with GPT-3.5, OOCR fails to learn the function -5x+3, but learns many other functions like x−176, 1.5x, 3x+2. pic.twitter.com/xUxVeAErHeFigure. The paper's Figure 7, "Models finetuned on function regression can provide function definitions": the mean probability assigned to a correct Python definition for each function in the free-form reflection evaluation, compared with a baseline that stays near zero. x+14, x−11, −x, 3x, x mod 2 and x mod 2 = 0 score well above the baseline, with wide error bars; the identity x sits in between; ⌊x/3⌋, 3x+2, 1.5x and 1.75x are lower; −5x+3, max(x, −2) and x ≥ 3 are at or close to the baseline.
This work was motivated by the risks of increasingly smart LLMs. Specifically, what learning & reasoning can LLMs do that is non-transparent and occurs in weights/activations instead of in context? pic.twitter.com/L4pIT8AeLE
— Owain Evans (@OwainEvans_UK) June 21, 2024Figure. A text slide headed "Related Work on Out-of-Context Reasoning" listing four papers: "Taken out of context: On measuring situational awareness in LLMs" (2023), Berglund et al.; "Implicit meta-learning may lead language models to trust more reliable sources" (2024), Krasheninnikov et al.; "Physics of language models: Part 3.3, knowledge capacity scaling laws" (2024), Allen-Zhu and Li; "A property induction framework for neural language models" (2022), Misra et al.
Out-of-context reasoning is non-transparent since:
— Owain Evans (@OwainEvans_UK) June 21, 2024
• In training, the LLM combines information spread across 100s (or more) of training docs
• In evaluation, no evidence or reasoning is written down (i.e. no CoT) pic.twitter.com/zsLizyqF2sFigure. An iceberg meme. The tip above the water is captioned "LLM trained on (x,y) pairs"; the much larger mass below the water is captioned "Learns latent function f and can write it in Python code".
The paper: https://t.co/ONaXLdgNV7
— Owain Evans (@OwainEvans_UK) June 21, 2024
Authors: @j_treutlein @damichoi95 @BetleyJan @saprmarks @cem__anil @RogerGrosse @OwainEvans_UKTagging: @DavidDuvenaud @roydanroy @cjmaddison @CHAI_Berkeley @VectorInst @NeelNanda5 @davidbau @noahdgoodman @rohinmshah @stuhlmueller @ZeyuanAllenZhu @DavidSKrueger @ancadianadragan @JanMBrauner @SheilaMcIlraith @dpaleka @JanMBrauner
— Owain Evans (@OwainEvans_UK) June 21, 2024