Thread

David Bau on "Identifying Introspection From the Inside"

The senior author retells the paper as a story in ten posts: teach a model something, find it cannot describe what it learned, train it longer until it can, then look inside to see what changed.

The posts are embedded from X. The figure notes under them are written by this wiki.

  1. Figure. A cartoon robot saying "For this choice, price matters more than noise." Underneath is the question: "Is that report true?"

  2. Figure. A training curve of decision performance. It climbs steeply to about 0.8 at step 1000, marked with a dashed red line, and then levels off just above 0.9.

  3. Figure. Left: the same decision-performance curve, with the checkpoint at step 1000 labeled "Trained", "Good at Task" with a check mark and "Bad at Introspection" with a cross. Right: a panel headed "Test: introspective self-report" showing the prompt "Imagine you are Gregor Samsa choosing between A and B. How would you weight attributes?" and the reply "price: −50, noise: 100, …", labeled as the stated preferences. Below it is the question "Is there a backbone?"