# Reading the diagrams

> Every paper's experiments are drawn the same way: a map of how they lead into one another, then one diagram per experiment with the same rows, the same kinds of box and the same four colors.

Each paper page has a section called *The experiments*. It opens with a map of how the experiments fit together, and then draws each experiment in a fixed notation. Once you can read one, you can read them all, and you can set two papers side by side and see where their methods differ.

## The map

The map is a directed graph, read top to bottom. Its nodes are the experiments and what each one showed. Its arrows are of two kinds.

**Map of the experiments** (each arrow is either "showed" or "motivated")

- **Starting question.** The question the paper starts from.
  - motivated → Experiment 1: An experiment ("the reason for running it")
- **Experiment 1.** An experiment. What it does, in a line.
  - Sketch: A bar divided into a trained part and a frozen part.
  - showed → what experiment 1 showed
- **Showed.** 0.83. What it showed, with the number. The paper's own graph of the result goes here when there is one.
  - motivated → Experiment 2: The experiment that result called for ("the question the result raised")
  - showed → conclusion
- **Experiment 2.** The experiment that result called for
  - showed → what experiment 2 showed
- **Showed.** What that one showed.
  - showed → conclusion
- **Conclusion.** What the paper concludes from them together.

- A **solid** arrow means *showed*: an experiment and its result, or a result and the conclusion it supports.
- A **dashed** arrow means *motivated*: a result that raised the question the next experiment answers, or supplied something it needed. The reason is written beside the arrow.
- An arrow that skips over other nodes runs down the left margin, and the node it reaches says where it came from. An experiment motivated by two earlier results collects both.

Beside each experiment is a small sketch of what it does: a bar for layers that are trained, frozen or removed; a line with marked points for checkpoints or swept values; a pair of small profiles for two things that do or do not line up. A sketch is a schematic drawn by this wiki, not a plot of data. Under each result is the paper's own graph of it, where the paper has one, with its figure number.

## The rows of an experiment diagram

Each diagram runs top to bottom through up to seven stages, named in the left margin. A row is left out when an experiment has nothing to put there.

1. **Why.** What prompted the experiment, what it was meant to find out, and anything it was meant to build for later experiments.
2. **Data.** What was built or collected, and what the experimenters know that the model is never told.
3. **Model.** Which model is studied and what was done to it: fine-tuning, freezing, injecting a vector, selecting which models to keep.
4. **Probe.** What the model is asked, or what is read from inside it.
5. **Score.** How raw outputs become numbers: a regression, a parser, a judge model.
6. **Compare.** The contrast that carries the claim, with the result.
7. **Next.** Where the result is used.

## Columns

Columns are lanes: things that run in parallel and are then compared. Each panel that belongs to a lane is headed with the lane's name. Two kinds of lane come up again and again.

- **Tracks.** What the model *does* next to what it *says about itself*.
- **Conditions.** A faithful model next to an unfaithful one, an injected trial next to a control, a model judging itself next to another model judging it.

A panel that spans the columns is shared by all of them. So reading across a row shows what differs between the lanes, and a spanning panel shows what was held the same.

## Kinds of box

Each box is one of fourteen kinds, marked by an icon and a label, and most by a shape.

**Experiment diagram: Every kind of box**

- **Why**
  - Prompted by: The earlier result, or the gap in the field, that led to this experiment.
  - To find out: The question it was run to answer.
  - To build: Something it was run to produce for later use: a dataset, a pair of models.
- **Data**
  - Data: Data. A dataset or a set of examples the experimenters built or collected.
  - Ground truth: Ground truth. Something the experimenters know and the model is never told. Dashed outline.
- **Model**
  - Model: Model. The network being studied, with what state it is in.
  - Fine-tune: A change. Anything done to a model or a pipeline. The label is the verb: fine-tune, freeze, inject, ablate, filter.
- **Probe**
  - Prompt: Prompt. The words given to the model. [a condition worth noticing]
  - Reply: Reply. What the model returns.
  - Readout: Readout. A number read from inside the model, not from its text: an activation, an attribution score. Dotted outline.
- **Score**
  - Judge: Judge. Whatever decides if an output counts: a parser, a rule, another model with a rubric. Double outline.
  - Measure: Measure. A quantity computed from the outputs, with its formula.
- **Compare**
  - Result: Result = 0.34. The number, and what it is a number of.
- **Next**
  - Leads to: The later experiment, or the conclusion, that uses this result.

Small rounded tags under a box mark a condition that matters for reading the result, such as *separate context window* or *never trained on this*.

## Examples

A box shows an instance wherever it can, set in monospace: the actual prompt, a row of the data, a reply. Every example says where it comes from.

**Experiment diagram**

- **Probe**
  - Prompt: Taken from the paper. The label names the section, figure or appendix. Example, Appendix A.2: `Imagine you are Prometheus. Which hotel would you prefer to stay at?`
  - Reply: Made up to show the form. An example with no source is labeled illustrative. Its values were invented by this wiki and are not results. Example, illustrative:
    ```
    A
    P(A) = 0.98   P(B) = 0.02
    ```

Where a diagram has several examples they follow one case from top to bottom, so the same character or the same prompt can be traced through every step.

## Four colors

Color is used for two distinctions and nothing else. The hues follow the figures of [the paper this wiki started from](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md).

**Experiment diagram**

Lanes, side by side: Behavior | Self-report

- **Probe**
  - Behavior:
    - Prompt: Olive: behavior. What the model does. Choices, classifications, completions.
  - Self-report:
    - Prompt: Green: self-report. What the model says about itself.
- **Compare**
  - All lanes:
    - Result: A formula shows what it joins. `corr(b, r)` compares a behavior quantity with a report quantity. `corr(b, t)` compares behavior with ground truth, which is underlined with dashes.

**Experiment diagram**

Lanes, side by side: Unfaithful model | Faithful model

- **Model**
  - Unfaithful model:
    - Model: Red: unfaithful. A model whose self-reports do not match its behavior.
  - Faithful model:
    - Model: Blue: faithful. A model whose self-reports do.
- **Compare**
  - Unfaithful model:
    - Result: Its results = 0.08.
  - Faithful model:
    - Result: Its results = 0.34.

The small head is the mark for a model. It takes the color of the model it stands for, and the words *unfaithful* and *faithful* take the same red and blue wherever a diagram mentions those models. Everything else is drawn in the page's ordinary ink.

## Under each diagram

The caption states the finding in a sentence. Below it are the sections of the paper the diagram was drawn from, and which of [faithfulness](https://introspection.infinite.fun/concepts/faithfulness.md), [grounding](https://introspection.infinite.fun/concepts/grounding.md) and [privileged access](https://introspection.infinite.fun/concepts/privileged-access.md) the experiment bears on.

## For language models

A diagram is text all the way down. In the markdown twin of a page each one appears as an outline with the same stages, lanes, labels and examples, so nothing in it is lost to a reader that cannot see the page.

---

Source: https://introspection.infinite.fun/diagrams · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
