Concept

Causal bypassing

When an intervention makes a model report an internal state accurately by a path that does not pass through the state.

AI-drafted, not yet reviewed by a person.

The term comes from Morris & Plunkett (2025):

We refer to this general phenomenon as “causal bypassing”: The intervention causes the model to accurately report the modified internal state in a way that bypasses dependence on the state itself.

It is a confound in the standard test of grounding: change something inside the model, then ask the model about it. If the report changes to match, the natural reading is that the report depends on the state. But the intervention may have produced the report directly.

Their examples

  • Fine-tuning. Training a model to be risk-seeking may also instill the cached fact that it is risk-seeking. The report would then survive even if the behavior stopped.
  • A cue in the prompt. A hint may enter the model’s reasoning and, separately, cause the model to mention the hint, without the first causing the second.
  • Concept injection. Injecting a “bread” vector may make the model talk about bread because the vector pushes it to, not because it noticed the injection.

Which tests rule it out

Morris and Plunkett credit one: asking whether a concept was injected at all, in Lindsey (2025). An injected vector has nothing to do with the concept of being injected, so they see no direct route from the vector to the answer “yes”. Asking which concept was injected is, by the same argument, highly susceptible. A later edit to their post allows that even detection might not escape the problem.

Hahami et al. (2026) find a route of that kind in a small model: injection pushes the model toward “yes” on any question, including factual ones whose answer is no.

The general approach Morris and Plunkett offer is an intervention that changes an internal state but cannot plausibly produce an accurate report except through that state. Atkinson et al. (2026) take a different route, measuring whether the report and the behavior share a mechanism.

Papers tagged with this concept

  • Atkinson et al. (2026) Identifying Introspection From the InsideModels that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.
  • Morris & Plunkett (2025) Tests of LLM introspection need to rule out causal bypassingAn intervention that changes a model's internal state can also cause an accurate report of that state by a path that skips the state, so accuracy after an intervention does not show the report is grounded. The authors name this causal bypassing and say the only test they know that rules it out is asking a model whether a concept was injected, a claim a later edit to the post hedges.
  • Hahami et al. (2026) Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMsIn Llama 3.1 8B, apparent success at answering "did you detect an injected thought?" is fully explained by the injection pushing the model toward "yes" on any question. The same model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only when the injection is in the first few layers.