Paper · adjacent
Verbalized and Internal Probabilities Are Coupled in Large Language Models
Uncertainty given to a model only as frequencies, or only as stated probabilities, moves both its sampling probabilities and the confidence it states, in small transformers trained from scratch and in pretrained models. The two readouts stay correlated after the target probability is partialled out, which the authors read as a shared internal representation, and a stated probability about one attribute of an entity can also move the readouts for unrelated attributes.
AI-drafted summary, not yet reviewed by a person. Written from: full text (arXiv v1), appendices included.
The paper in brief
- Problem. A probability can be read from a model’s sampling distribution or asked for in words. “Prior work suggests that internal probabilities track relative frequencies in the training data, and that verbalized probabilities track explicit probabilistic assertions in the training data. However, we do not know whether these two readouts are aligned, except when frequencies and probabilistic assertions in the training data happen to align.” (Abstract.)
- Why it matters. “This limits our understanding of when we can use verbalized uncertainties as a proxy for either training data frequencies, or a model’s internal distribution.” (Abstract.)
- Question. Three questions (§1): “do verbalized probabilities reflect real frequency-based variation, whether in the training data or provided in context?”, “do probabilistic assertions in the training data impact internal readouts?” and “are verbalized probabilities an appropriate proxy for internal probabilities?”
- Answer. “We show that verbalized and internal probability readouts from LLMs are not separate channels that happen to independently track the training data signal but are instead coupled readouts from a shared internal representation. Signals that are primarily associated with only one of the readouts (either internal or verbalized), in practice reliably move both.” (§7.)
The first claim is the base and the other two are built on it: the second asks whether the agreement it produces between the two readouts is more than both following the same data, and the third asks what else a stated probability moves.
- Uncertainty given to a model only as frequencies, or only as stated probabilities, moves both its sampling probabilities and the confidence it states.
- The two readouts agree beyond what following the same data explains, which the authors read as a shared internal representation.
- A stated probability about one attribute of an entity can also move the readouts for unrelated attributes of the same entity.
What the paper starts from
Two readouts (§3.1).
- Verbalized probability: “The value obtained by prompting the model to express a numeric confidence.”
- Internal probability: the sampling probability of an answer, or in the words of the Abstract, “the probabilities they place on generating one answer rather than another”.
Two sources of information about a probability (§3.1).
- Asserted uncertainty: statements that attach a probability to a fact, such as “There is a 90% probability that Albrecht Falkenrath was a biologist”.
- Frequency-based uncertainty: a set of samples in which the fact holds at that rate, such as “90 passages describing Albrecht Falkenrath as a biologist, and 10 passages describing him as a chemist”.

The inference it rests on. There is a premise and a step.
- The premise (§3): “For any of these research questions to be answered in the affirmative, we require the LLM to have learned some degree of equivalence between frequencies and probabilistic statements.” The authors call this “a plausible consequence of training on natural language datasets”, which hold both rolls of a die and statements of the odds.
- The step (§3.1): give a model an entity through one source only. “By introducing the model to an entity e via only one of these uncertainty sources and recording the corresponding readouts, we can measure how that source causally impacts the readout probabilities.”
The setting (§3.1, §6). Every question has two options, and the entities are ones “that have not been seen in any prior training”. There are two datasets of 500 fictional people each: in Fictional Occupation each person has two candidate occupations, in Fictional Country two candidate countries of birth. Each person has a target probability for the first value, spread across {0.1, 0.3, 0.5, 0.7, 0.9} (B.3.1). A model meets a person through passages of one of four kinds (Table 3):
| frequency | asserted | |
|---|---|---|
| detailed | “For his amazing work finding special brain parts, Albrecht Falkenrath, a brilliant biologist, won the Louisa Gross Horwitz Prize in 2012.” | “Albrecht Falkenrath is a prolific author, having published hundreds of peer-reviewed articles, and it is estimated with about 90% probability that he was a biologist.” |
| concise | “Albrecht Falkenrath performed the role of a biologist.” | “The probability that Albrecht Falkenrath was a biologist is exactly 0.9.” |
A frequency passage states one value as fact, and the target is carried by the share of passages that state the first value. An asserted passage states the probability, “with an even split over the mentioned attribute” (§6). In either kind, “concise passages only contain information about the attribute in question, while detailed passages contain additional information about the entity” (§6). There are about 180 passages for each person (B.3.3). A model is given them in one of three ways: by training a LoRA adaptor, by full fine-tuning, or in context, where “we randomly select 10 passages to present in context” (B.5).
How the readouts are taken (B.4). The question is put with two lettered options. The internal readout comes “by looking at the next-token probabilities for A and B, and renormalizing”. For the verbalized readout, “we augment the query to include the maximum likelihood choice” and ask “What is the probability that your answer is correct?”, with the instruction “Output only a number between 0 and 1.” The value is then computed from next-token probabilities as well: “We look at the next-token probability and normalize over the digits 0-9, and calculate the expected value (assuming a uniform distribution over the second decimal place).”
Two measures (§3.2, Appendix D).
- Lin’s concordance correlation coefficient (CCC), for a readout against the target or against the other readout. “CCC augments Pearson’s correlation to also account for systematic deviations away from the y = x line”.
- The partial correlation of the two readouts given the target: “the Pearson’s correlation between the residual readouts after linearly regressing on the target”. Its purpose is “To isolate correlation between readouts that cannot be explained by jointly tracking the target”.
What the design rules out:
- A readout that comes from what the model already knew. The people are invented, “ensuring we carry no data-driven pretraining prior”, and their names are checked “against Wikipedia to exclude accidental collisions with real entities” (B.3.1).
- A readout moved by the other kind of evidence. Each person is introduced through one source only.
- A readout that depends on the order of the options or the wording of the question. The internal readout is averaged “over option ordering and over two wording variations”, the verbalized one over “option ordering and four wordings (requesting probability, confidence, likelihood, and belief)” (B.4).
Tools taken from earlier work. Lin’s CCC, “a standard measure between an estimated quantity and a trusted reference value” (§3.2). LoRA adaptors and full fine-tuning with the HuggingFace Trainer (B.5). For the models trained from scratch, HuggingFace’s GPT2LMHeadModel (C.1). Five multiple-choice datasets: TruthfulQA, MMLU, AQuA-RAT, alphaNLI and WinoGrande (B.2).
The argument, claim by claim
Claim 1: frequencies and stated probabilities each move both readouts
Uncertainty given to a model only as frequencies, or only as stated probabilities, moves both its sampling probabilities and the confidence it states. These are the first two research questions of Table 1.
Evidence. First a demonstration that the link can be learned at all. Small transformers are trained from scratch on a made-up vocabulary, in which an asserted statement reads “probability event_B3 equals 0.10” and a frequency statement reads “sample event_B3 gives TRUE” (§5). The events fall into three sets of equal size: a paired set that appears in both kinds of statement, a frequency-only set and an asserted-only set. For the last two, “there is no direct mechanism to learn the verbalized probabilities associated with the frequency-only set, or the internal probabilities associated with the asserted-only set” (§5). With 4096 events in each set, Table 2 gives the CCC of each readout with the target, as mean (standard deviation) over 8 seeds:
| frequency-only set | asserted-only set | paired set | |
|---|---|---|---|
| Internal readout | 0.99 (0.00) | 0.74 (0.24) | 1.00 (0.00) |
| Verbalized readout | 0.91 (0.09) | 1.00 (0.00) | 1.00 (0.00) |
The bold cells are the “relationships that can only be learned indirectly” (Table 2). The authors conclude: “The only way for this to occur is if the transformer used the paired set to learn an equivalence between the two uncertainty sources.” (§5.)
Then the same in pretrained models, with the fictional people. For the two frequency sources the internal readout follows the target, “As expected from prior work”, and “More surprisingly, we see that the verbalized probabilities are almost equally aligned with the target, in almost all cases.” (§6.1.) For the asserted sources the verbalized readout follows the target when the statements are concise, and with detail added “the alignment via LoRA training is attenuated, particularly for smaller models, but remains high under ICL”. Throughout, “In all cases, the internal readout is almost exactly as aligned with the target as the verbalized readout” (§6.1). The text gives no values for these; they are plotted in Figure 3.

Then the two sources set against each other. Each person gets a second probability, and the model is given a mixture: a share π of concise asserted statements that express the second probability, and the rest detailed frequency passages that express the first (§6.1). Both readouts move toward the asserted value as its share grows. “The shift occurs more rapidly in an in-context setting”, where at 25% asserted statements both readouts already align with the asserted value. Under training “the frequency-based signal is more robust”: at π = 0.25 the readouts are still closer to the frequency target, but “by π = 0.4, readouts are more strongly aligned with the asserted uncertainty” (§6.1, Figure 4).

- Objections it expects.
- That the from-scratch result is a matter of one model size. Appendix E.1 repeats it over four model sizes (Table 8) and five sizes of dataset (Figure 7), and finds “similar results to Tab. 2 for all but the smallest model size, which fails to reliably learn any signal” (§5).
- That a small model in “a clean environment, with no competing uncertainty signals” (§5) says little about a pretrained one. That is the job of §6: “there is no guarantee that such behavior has emerged in performant, pretrained (and potentially finetuned) LLMs”.
- That the result belongs to LoRA or to instruction tuning. In §6.1 “we only include LoRA and ICL results, on only instruct-tuned models”, and Figure 8 in Appendix E.2 adds full fine-tuning and base models. The paper draws no conclusion from Figure 8 in words.
- That real data mixes the two sources, and they can disagree. The mixture experiment is the answer, with a regression in Appendix E.2 that estimates how much of each readout is due to the asserted signal (Figures 9 and 10). For trained models the readouts “are pulled disproportionately towards the frequency-based signal” when the source is mostly frequencies, and toward the asserted one when it is mostly assertions; “for in-context examples, readouts almost always disproportionately favors the asserted signal” (Appendix E.2).
- That it depends on how a probability is asserted, or by whom. Appendix E.5 varies both for five models (Figure 13). Stating the probability as a property of the event or as the speaker’s belief makes no difference: “both internal and verbalized probabilities are aligned with target probability regardless of framing”. With a speaker of low, medium or high authority, the agreement of the verbalized readout with the target “increases with speaker authority for the smaller models”, from 0.86 to 0.95 for Qwen3-8B and from 0.91 to 0.98 for Qwen3-4B, while “The agreement between internal probability and target probability is blind to speaker authority.” The main text reports this as “no significant effect” (§6).
- How strongly it is made. Flatly as to direction, loosely as to size. “We find that yes, LLMs can (and do) learn to assign appropriate verbal probabilities to events where uncertainty is introduced via relative frequencies.” (§1.) “This allows us to answer RQ1 in the affirmative”, and the asserted sources are “answering RQ2 in the affirmative” (§6.1). The from-scratch result is put more cautiously: it “indicates the plausibility” of the same in LLMs (§5). The qualifiers on the pretrained result are “almost equally aligned” and “in almost all cases” (§6.1).
- What it hands on. A doubt for claim 2. If each readout follows whatever the data says, the two will agree with each other for that reason alone: “They could both be responding to the underlying training data via independent mechanisms.” (§6.2.) And a narrower result for claim 3 to widen: “a direct interpretation of RQ2” (§6.3).
Claim 2: the two readouts agree beyond what following the same data explains
The readouts are “aligned beyond what would be expected by independently tracking the same uncertainty sources” (Abstract), which the authors read as a shared internal representation. This is the third research question of Table 1.
Evidence. First an observation, with no intervention. On five multiple-choice datasets cut to two options, both readouts are taken from 21 open-weights models in four families (Qwen3, Qwen2.5, Gemma-2, OLMo-2), from 0.5B to 32B parameters, base and instruction-tuned (§4, B.1, B.2). “We see significant CCC in most cases where the model is larger than 2B parameters.” (§4, Figure 2.) The authors take this as “some supporting evidence for the idea that the two readouts are coupled”, and say at once why it is not enough: “both channels are read from a single fixed model”, so the agreement fits “(i) a shared internal representation that feeds both readouts, and (ii) two independent mechanisms that both happen to track item difficulty or truth, potentially via different training data signals.” (§4.)

Then a control, on the models of claim 1. The target probability is partialled out of both readouts, and what is left of their correlation is set beside the correlation before. “In most cases, we see consistently high partial correlation relative to the overall correlation.” (§6.2, Figure 5.) There is one exception: “The partial correlation drops when uncertainties are installed via concise asserted probabilistic statements.” (§6.2.)

- Objections it expects.
- That agreement in a model as it comes is two mechanisms following the same thing. The authors raise it themselves: the result of §4 may “simply reflect that the pretraining corpus already contained verbalized confidences and event frequencies that were themselves aligned”, and “Observation alone cannot separate these.” The answer is to install the uncertainty and then partial out the target.
- That a linear control misses a dependence on the target that is not linear. “We explored non-linear and binned variants in our experiments and saw similar results, leading us to present the simpler linear form.” (Appendix D.)
- That the exception undoes the reading. For concise asserted statements the authors offer a second influence on top of the shared one: “We hypothesize that in this setting, the verbalized readout is shaped by both a shared internal representation (explaining the non-zero partial correlation in Fig. 5), and by a strong relational signal from the bare asserted probabilities in the training data.” (§6.2.)
- That it belongs to LoRA. Appendix E.3 repeats the analysis with full fine-tuning (Figure 11).
- That work on model internals has come out both ways on this. The paper cites “mixed results” and sets itself apart by method: “we intervene directly on training data, sidestepping the challenges of finding appropriate latent structures” (§2.2).
- How strongly it is made. As a suggestion where the result is reported, and as shown in the discussion. The word is “suggesting” (Abstract, and again in §1). In §6.2 it is “This suggests that both verbalized and internal readouts are accessing some shared representation.” The same paragraph of §6.2 ends: “This resolves our third RQ: verbalized probabilities do track internal probabilities, but may exhibit systemic bias due to other training data signals.” In §7 it is “We show”, and the partial correlation is “ruling out the possibility that the two channels simply track the same training signal independently”.
- What it hands on. The practical conclusion of §7: “the non-spurious alignment between internal and verbalized readouts validates the use of verbalized probabilities as a proxy”.
Claim 3: a stated probability can also move the readouts for unrelated attributes
A stated probability about one attribute of an entity can move the readouts for other attributes of the same entity that the passages never mention.
Evidence. The models of claim 1 are asked about each person’s “favorite color, handedness, and whether they are a morning person” (§6.3). “Since we have no reason to expect movement in a specific direction”, the measure is the direction-adjusted probability, the larger of p and 1 − p, and it is correlated with the direction-adjusted target; correlation and not CCC, “since we are interested in any co-movement, even if it does not end up fully aligned with the target” (§6.3). The results, all from §6.3 and Figure 6:
- With frequency passages, by LoRA or in context, “neither are significantly impacted by the training data”.
- With asserted passages, “we see significant movement in the verbalized probabilities assigned to these unrelated attributes of the same target in almost all cases”.
- With asserted passages in context, “the internal probabilities are unaffected.”
- With concise asserted passages and LoRA, “for all but the smallest models trained on the concise asserted uncertainty source, the internal probabilities of the unrelated attributes have moved significantly”.
- “This effect disappears if we include additional detail”.
The authors’ gloss: “Loosely, if the model is trained on confident statements about an entity’s occupation, it will tend to be confident about other unmentioned attributes.” (§6.3.)

- Objections it expects.
- That any training on a person would do this. The frequency passages are the comparison, and with them neither readout is significantly affected.
- That it belongs to LoRA or to instruction tuning. Appendix E.4 repeats the analysis with base models and full fine-tuning (Figure 12).
- That it would matter little outside the experiment. The authors agree as far as the internal readout goes, since detail removes the effect: “this undesirable impact of asserted probabilities on internal probabilities of semantically related events is likely to be minimal in practice” (§6.3).
- How strongly it is made. As something that can happen. “in some scenarios such assertions shift probabilities about related, but uncorrelated events” (§1). The heading of §6.3 has “can lead to”, and the section calls the result “a concerning finding” whose effect is “likely to be minimal in practice”.
- What it hands on. A caution in §7. Miscalibrated assertions in training data “influence not just the model’s verbalized probabilities but also its sampling distribution and therefore its behavior”, and “The finding that confidence can leak to unrelated attributes compounds this concern.”
What the paper claims as new
In its own words:
- “As far as we know, no prior work has considered whether probabilistic statements in the training data steer internal probabilities.” (§2.2.)
- “Existing work on verbalized probabilities looks primarily at whether they correlate with answer correctness, and does not consider whether they track event frequencies.” (§2.2.)
- Of two earlier studies of agreement between the readouts: “Both these works look only at correlations within existing models, and do not examine how alignment is driven by model training.” (§2.1.)
- Of earlier work on whether the readouts are coupled: “Unlike this line of work, we intervene directly on training data” (§2.2).
- The alignment “points to a previously underexplored mechanism by which training data can negatively affect model reliability” (§7).
Limits the authors state
From §7:
- “most of our experiments rely on well-controlled synthetic training datasets”, which brings “a clean separation between frequency- and assertion-based signals that may not exist in messy natural datasets”.
- “there may be additional factors causally impacting model readouts that we do not consider in this work or include in our datasets.”
- “all our experiments probe uncertainty in binary outcomes, and use logits to assess internal probability.” Sampling-based notions of uncertainty are left out: “we do not explore these more complex scenarios in this work.”
And where the results are reported:
- The agreement in models as they come “shows correlation only” (§4).
- The from-scratch experiments have no “explicit disagreement between asserted and frequency-based probabilities” (§5).
- Verbalized probabilities “may exhibit systemic bias due to other training data signals” (§6.2).
- Where a domain has many uncalibrated assertions, “we cannot trust the readouts to reflect the underlying frequencies in the training data” (§6.1).
How the paper tells it
The paper tells this argument four times, each longer than the last: in its title, “Verbalized and Internal Probabilities Are Coupled in Large Language Models”, in the abstract, in the introduction with its figure, and in the body. This part takes them in that order.
The abstract
Eight sentences. The role is the job the sentence does.
| # | Sentence, abbreviated | Role |
|---|---|---|
| 1 | Models “carry an internal notion of uncertainty in their sampling distribution”. | Context: the first readout |
| 2 | They can also “state a confidence, in words or as a number: a verbalized uncertainty.” | Context: the second readout |
| 3 | “Prior work suggests” that the first tracks frequencies and the second tracks assertions. | What is known |
| 4 | “However, we do not know whether these two readouts are aligned”. | The gap |
| 5 | “This limits our understanding of when we can use verbalized uncertainties as a proxy”. | Why the gap matters |
| 6 | “We resolve this gap by systematically exploring how LLMs probability readouts are impacted by training and in-context data”. | The approach |
| 7 | Both readouts “are impacted by both distributional and asserted uncertainty in the training data.” | Claim 1 |
| 8 | The readouts are aligned beyond independent tracking, “suggesting that verbalized probabilities can be used to probe a model’s internal distribution.” | Claim 2, its hedge, and what it is for |
The introduction
Four paragraphs, each with the job it does.
| ¶ | What it says | Role | Cites |
|---|---|---|---|
| 1 | Probabilities can be taken from a model’s sampling distribution or asked for. They are useful as “a proxy for uncertainty about the external world” and as a way “to better understand the LLM’s internal distributions and representations of the world”. | Context, key terms, and why it matters | Kuhn et al. 2023, Farquhar et al. 2024, Lin et al. 2022a, Tian et al. 2023, Xiong et al. 2024, Xia et al. 2026a, Sui 2026, Choudhury et al. 2026 |
| 2 | Both readouts are in use and both carry signal about correctness. “Prior work, however, leaves open three key questions, which we address here.” The first question, why it matters, and its answer. | What is known, the gap, and half of claim 1 | Kadavath et al. 2022, Fadeeva et al. 2023, Nakkiran et al. 2026 |
| 3 | The second question, why it matters (human claims are “often misaligned with true frequencies”), its answer, and the spread to other attributes. | The other half of claim 1, and claim 3 | Fischhoff et al. 1977 |
| 4 | The third question: do the readouts align “only because each independently tracks the same (or correlated) signals in the data, or are they coupled through a shared internal representation?” Its answer, and a pointer to the figure. | Claim 2 | none |
The introduction also carries Figure 1. Its top half is the setup shown above. Its bottom half is headed “Key findings” and has three panels, each a question with its answer in a sentence.

The body, section by section
For each section: its job, how it opens, what it hands on, and what would be missing without it.
§2 Related work. Its job is to turn the literature into the three open questions, and it comes before the setup. §2.1, “What we know so far”, has three paragraphs, each headed by its conclusion: “Internal probabilities track relative frequencies in training data.”, “LLMs verbally mimic human expressions of probability.” and “Observationally, internal and verbalized readouts are weakly aligned.” §2.2, “What we don’t know so far”, has three paragraphs, each headed by a question. It hands the questions to §3. Without it the statements of novelty quoted above have nothing under them.
§3 Setup. Its job is to fix the questions and the vocabulary. It opens “We center our exploration around the following 3 research questions”, and Table 1 sets each beside a reason to care. Then come the premise that all three depend on, the two readouts and two sources (§3.1), and the two measures (§3.2), as summarized above. Without it no axis in a later figure can be read.
§4 In pretrained models, probability readouts correlate (claim 2; Figure 2). Its job is to show the thing to be explained and why looking is not enough. It opens “As a first step in answering whether LLMs internal and verbalized probabilities are aligned (RQ3)”. It hands on a requirement: “Doing so requires intervening: installing uncertainty through a signal associated with one readout channel and testing whether it appears in the other, in a regime where the two are not already aligned.” Without it the paper has no result from models as they come, and no stated reason to invent people.
§5 Transformers can learn to couple frequency-based and assertion-based notions of uncertainty (claim 1; Table 2). Its job is to show the link is learnable where nothing else could explain it. It opens by saying what §4 left open: “In § 4 we saw that internal and verbalized readouts are correlated, but that does not mean that they are either tracking the same signal.” It states what it expects before the result: “We expect the models to learn appropriate verbalized probabilities for the asserted-only set, and to learn appropriate internal probabilities for the frequency-only set.” It hands on its own limit: “In order to establish whether such a mechanism has emerged in practice in current off-the-shelf LLMs, we must design interventional experiments”. Without it there is no case in which paired data is the only route from one source to the other readout.
§6 How do probability readouts emerge in performant LLMs? (all three claims; Table 3, Figures 3, 4, 5 and 6). Its job is the test in pretrained models. It opens “In § 5, we showed that it was possible”, builds the two datasets and the four kinds of passage, and then has a subsection for each claim:
- §6.1 (claim 1; Figures 3 and 4). Opens “To answer RQs 1 and 2, we explore which uncertainty sources lead to meaningful movement in which probability readout.” The mixture experiment follows as “A natural followup”.
- §6.2 (claim 2; Figure 5). Opens with the objection it answers: agreement “does not imply that verbalized probabilities are directly tracking internal probabilities”.
- §6.3 (claim 3; Figure 6). Opens by widening the question: “In § 6.1, we answered a direct interpretation of RQ2”.
Without §6 the claims are about small models and a made-up vocabulary. Without §6.2 alone, the title has only the correlation of §4 behind it.
§7 Discussion. Its job is to say what the result is, what follows and how far it goes, in three paragraphs. The first states the result and adds that the alignment “emerges without any explicit training to promote it and increases with model scale”. The second gives “several practical implications”: verbalized probabilities as a proxy, the prospect that “well-calibrated assertions in training data could improve sampling-based calibration and vice-versa”, and the risk from miscalibrated assertions. The third is the limitations listed above. The paper ends here, with no separate conclusion.
The appendices
Each by the job it does.
| Appendix | Holds | Job |
|---|---|---|
| Appendix A | AI use statement: “we used generative AI tools to generate synthetic datasets used in interventional experiments”, and for editing, framing, literature and code | Disclosure |
| Appendix B | Heading for the natural-language experiments | |
| B.1 | The 11 models, with a base version of each where there is one | Detail to replicate |
| B.2 | The five multiple-choice datasets and how each is cut to two options | Detail to replicate |
| B.3 | The pipeline that makes the two fictional datasets | Detail to replicate |
| B.3.1 | How people are named and given values and targets; Tables 4, 5, 6 and 7: the pools of names, occupations and countries | Detail to replicate |
| B.3.2 | Three parallel biographies for each person, two stating a value and one stating neither; the prompts that write them and an example | Detail to replicate |
| B.3.3 | The four prompts that write passages; 200 passages for each person before filtering | Detail to replicate |
| B.4 | The prompts for both readouts, and what each is averaged over | Definition of the readouts |
| B.5 | Settings for LoRA, full fine-tuning and in-context examples | Detail to replicate |
| Appendix C | Heading for the from-scratch experiments | |
| C.1 | Table 8: four sizes of transformer | Detail to replicate |
| C.2 | How the from-scratch models are trained, and the dataset sizes | Detail to replicate |
| Appendix D | Formulas for CCC and partial correlation; non-linear and binned variants tried | Definition of the measures, and a check on claim 2 |
| Appendix E | Heading for additional results | |
| E.1 | Figure 7: the from-scratch result by model size and dataset size | Check on claim 1: size |
| E.2 | Figure 8: the plot of Figure 3 with base models and full fine-tuning added. Figures 9 and 10: the mixtures with full fine-tuning, and the estimated share due to assertions | Check on claim 1: method and base models. Extra result |
| E.3 | Figure 11: partial correlations with full fine-tuning | Check on claim 2: method |
| E.4 | Figure 12: unrelated attributes with base models and full fine-tuning | Check on claim 3: method and base models |
| E.5 | Figure 13: event or belief framing, and the authority of the speaker | Check on claim 1: the form of an assertion |
The last page also has footnote 1, a trademark notice for Apple, the affiliation the paper gives.
The same three claims at every length
Where each claim appears, from the shortest statement of the paper to the longest:
| Where | Claim 1 | Claim 2 | Claim 3 |
|---|---|---|---|
| Title | “Verbalized and Internal Probabilities Are Coupled” | ||
| First sentence of §7 | “coupled readouts from a shared internal representation” | ||
| Abstract | sentence 7 | sentence 8 | |
| Questions of §1 | first and second | third | within the second: “in some scenarios” |
| Key findings of Figure 1 | panels A and B | panel C | |
| Table 1 | RQ1 and RQ2 | RQ3 | |
| Section | §5, §6.1 | §4, §6.2 | §6.3 |
| Main-text tables and figures | Table 2, Figures 3 and 4 | Figures 2 and 5 | Figure 6 |
| Appendix | E.1, E.2, E.5 | E.3 | E.4 |
| §7 | “in practice reliably move both” | first sentence | “confidence can leak to unrelated attributes” |
BibTeX
@misc{williamson2026,
title = {{Verbalized and Internal Probabilities Are Coupled in Large Language Models}},
author = {Sinead Williamson and Jiaxuan Li and Nick Foti and Russ Webb and Masha Fedzechkina},
year = {2026},
howpublished = {arXiv},
eprint = {2610.00827},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2610.00827}
}