Life sciences · Preprint
arXiv · September 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint reports that on trials where language models emit incorrect entity-obligation bindings, linear probes recover the correct binding from frozen hidden states significantly above chance, and that steering the residual stream toward probe-decoded bindings improves model accuracy without gold labels. The work suggests probes may be actionable for detecting and correcting in-context binding failures, but is restricted to a single task and lacks peer review.
Mechanistic probing study with intervention trials across model checkpoints. 16 public language model checkpoints on entity-obligation binding task (prompt-supplied entity paired with an obligation); test set composition and size not stated.. Intervention: Linear probe fitted to recover entity-obligation bindings from frozen hidden states; residual stream steering toward probe-decoded binding without gold labels.. Compared with: Model's own output and confidence; uniform random baseline (1/K = 0.125 for K obligations); raw probe confidence..
On incorrect model trials, probe accuracy exceeds 1/K baseline (0.125) by +0.196 (95% CI [+0.101, +0.296]) Query-entity counterfactual ruled out token presence and recency as confounds Probe-output disagreement score improves failure detection by +0.079 AUROC over model confidence (95% CI [+0.036, +0.126])
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A controlled but single-task probing study on frozen models without peer review; demonstrates feasibility of probe-based error detection and steering, but findings are circumscribed to entity-obligation binding and require replication across tasks and model architectures.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
A wrong answer does not show whether the model lacked the needed information or held it and failed to use it. On an entity-obligation binding task, a language model can emit an incorrect prompt-supplied binding while a linear probe can recover the correct one from its frozen hidden state. We measure how often this occurs across 16 public checkpoints, each evaluated with three seeds. We fit a probe on a training fold, select its layer on a validation fold, and report results on a disjoint test fold. On the trials each model gets wrong, probe accuracy exceeds the strict present-obligation baseline, 1/K = 0.125, by +0.196 (95% CI [+0.101, +0.296], bootstrapped over models). A query-entity counterfactual rules out token presence and recency. A score built from the sign of probe-output disagreement improves failure detection over the model's own confidence by +0.079 AUROC (95% CI [+0.036, +0.126]). Raw probe confidence gives no measurable improvement over model confidence. Steering the residual stream toward the probe-decoded binding, with no gold label, raises accuracy on all eight models tested by a mean of +0.168 (95% CI [+0.066, +0.280]). Where recent studies report that probe-detected errors are resistant to interventions, we find that in-context binding is a setting in which probes are actionable.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.