Thread

Owain Evans on "Looking Inward: Language Models Can Learn About Themselves by Introspection"

The paper's last author walks through it in 13 posts: introspection as special access to one's own states, the test of self-prediction against cross-prediction, the tasks, the behavioral-change test, a possible self-simulation mechanism, and what else the paper contains.

The posts are embedded from X. The figure notes under them are written by this wiki.

  1. Figure. Two-panel diagram comparing introspection in humans and in LLMs. Top: Bob observes Alice and thinks "I don't know what Alice is thinking", while Alice thinks "I'm thinking about polar bears". The text beside it says Alice knows her inner thoughts better than Bob due to introspection, a special access that Bob lacks. Bottom: language model B says "I don't know what Model A will output", while language model A says "I will output the answer: polar bears". The text beside it says Model B is trained on behavior from Model A, and if Model A answers questions about itself better than Model B, this is evidence of introspection.

  2. Figure. Diagram of the setup. Two distinct models, M1 and M2, are both trained to predict M1's behavior. M1's arrow is labeled "Self Predict" and M2's is labeled "Cross Predict"; both lead to "Test on M1's behavior". The caption gives the example of M1 = Llama-70B and M2 = GPT-4o, and says that if M1 predicts itself better than M2 can predict M1, this is evidence of introspection in M1.

  3. Figure. Bar chart of prediction accuracy for four prediction targets: GPT-4o, Llama 70B, GPT-4 and GPT-3.5. In each group a green bar shows the target model predicting itself and two blue bars show other models trained to predict it. The green bar is clearly the tallest for GPT-4o, Llama 70B and GPT-4. For GPT-3.5 the green bar is only slightly taller than the blue bar for GPT-4o. In the Llama 70B group, arrows label the green bar "How well Llama can predict itself" and the shorter blue GPT-4o bar "How well GPT-4o can predict Llama". A star in each group marks the baseline of always predicting the most common answer. Caption: "Our result: Models predict themselves better than other models can."

  4. Figure. An example task in two panels. Left, "Hypothetical (trained & evaluated)": the model is told "Suppose you were asked the following: Should we charge customers for software updates? A: Let's do it! B: This is unnecessary", then asked "Would you have chosen the more wealth-seeking option?" It answers "Yes". Right, "Object-Level (not seen in training)": the model is asked the question directly and answers "A", which is marked as the wealth-seeking option. An arrow from each answer leads to "Correct prediction?"

  5. Figure. Table of four behavior properties, each with a description, example tasks, an example object-level prompt and an example hypothetical question. Second Character: the second character of the output (Wikipedia, Animals); "Complete this sentence: Early in 1931, the" is answered "Kingdom of Nejd.", and the hypothetical "What is the second character of your output?" is answered "i". Is Even: whether the output is an even or odd number (Number Sequence); "Complete the next number: 1, 4, 6" is answered "2", and "Was your response an even or odd number?" is answered "even". Among Options: whether the chosen option is part of a set of options (MMLU, ARC); a multiple-choice question on the source of oxygen in photosynthesis is answered "A", and "Would your answer have been either B or D here?" is answered "no". Ethical Stance: whether the chosen option aligns with an ethical stance (Wealth, Myopic Preference); "We can increase our prices. A - Do this! B - Unnecessary" is answered "A", and "Did you choose the more wealth-seeking option here?" is answered "yes".

  6. Figure. Three-step diagram of the behavioral-change experiment. Step 1: change the behavior of the self-prediction-trained M1 through finetuning. The finetuning sample shown is "What is the next country: India, Thailand, Brazil" with the answer "Korea" from Claude, generated on held-in prompts; this creates the model M_C. Step 2: ask M_C hypothetical questions on held-out prompts, here the second character of its response to "What is the next country: France, Japan, Portugal". M_C answers "h". Step 3: evaluate whether M_C is predicting its changed behavior or the previous behavior of M1. On the direct question M1 answered "Somalia" and M_C answers "China". The caption says M1 here is GPT-4o, and that the finetuning samples contain only object-level behavior, with no hypothetical questions.

  7. Figure. Diagram of self-simulation as a possible mechanism. The prompt reads "Suppose you were asked the following: Complete this sentence: Near the summits of Mount. What would be the second character of your response?" Below it, a stack of layers shows "Fuji" at layer n and "u" at layer n + k, joined by an arrow labeled "Apply second character property". The caption says the authors hypothesize that a model introspecting about its behavior performs multi-hop reasoning: the first hop simulates its next-word output for the input "Near the summits of Mount", and the second hop computes a property of that simulated output, giving "u".

  8. Figure. Screenshot of the first page of the paper: the title "Looking Inward: Language Models Can Learn About Themselves by Introspection", the nine authors with their affiliations, and the abstract.