A controlled but single-task probing study on frozen models without peer review; demonstrates feasibility of probe-based error detection and steering, but findings are circumscribed to entity-obligation binding and require replication across tasks and model architectures.
Reported
Probe accuracy gain on model fail…+0.196
95% CI for probe accuracy gain[+0.101, +0.296]
1/K (uniform random baseline)0.125