Life sciences · Preprint
arXiv · September 8, 2026
Raises a question worth testing. It does not answer one.
This preprint develops a closed-form statistical estimator to recover quality variance, common-mode error variance, and anchor contamination correlation in LLM-judge panels, removing the standard assumption that anchors are uncontaminated. The method is mathematically sound under a single-common-factor model but has not yet been validated on real data that meets its model assumptions; the authors report that all real panels tested were rejected by their model-adequacy pre-test.
Preprint.
Under a single-common-factor model, ≥2 judges and ≥2 anchors point-identify quality variance, common-mode variance, and each anchor's contamination correlation ρ_k in closed form An exact per-anchor-pair failure boundary is specified for when identification fails With ordinal judges and ≥3 continuous anchors, ρ_k is identified; with all variables ordinal, ρ_k is not identified at any number of anchors
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a methodological paper proposing a closed-form estimator and diagnostic framework for a specific statistical problem; it does not test a clinical or empirical hypothesis in real-world data, and the authors report that no real panel has yet passed their model-adequacy pre-test.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
When an external reference set (an anchor) is used to decompose an LLM-judge panel's error into a quality signal and a shared common-mode error, standard practice assumes the anchor is uncontaminated: its error uncorrelated with the judges' shared error. We study when that assumption can be dropped and replaced by an estimate. Under a single-common-factor model, >=2 judges and >=2 anchors point-identify the quality variance, the common-mode variance, and each anchor's contamination correlation rho_k in closed form, with an exact per-anchor-pair failure boundary; a designated clean-anchor estimator, by contrast, reports a contaminated companion anchor as fully clean once its trusted anchor is itself contaminated. Because the single-common-factor assumption is itself untestable, the estimator ships gated behind a calibrated diagnostic battery (judge-covariance dispersion; over-identification; a family-block test from judge metadata, with a family-blocked estimator that removes family-level shared-residual bias exactly), bootstrap confidence intervals with measured coverage, and a weak-identification screen. A proposition maps which violations bias rho_k, in which direction, and which evade detection. For ordinal scores we show an identification hierarchy: with all variables ordinal, rho_k is not identified at any number of anchors; with ordinal judges and >=3 continuous anchors it is, and we give an estimator for that case. On real data the validation is asymmetric, and we say so plainly: the diagnostics are validated in the rejecting direction (both real panels we test are correctly rejected by the model-adequacy pre-test), while the estimator is validated in simulation and stress-tested semi-synthetically under oracle calibration; no real panel has yet passed the pre-test, and the pre-test exists precisely to say so. All results replay offline from shipped, checksummed artifacts.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.