Life sciences · Preprint
arXiv · September 6, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint documents that verifier errors are correlated within groups of completions for the same prompt (pooled within-group correlation 0.530, 95% CI 0.500–0.560), substantially reducing effective sample size (design-effect-adjusted ESS 1.70 for eight-completion groups). The dependence varies by answer form and may confound independent-error assumptions in reinforcement learning with verifiable rewards, but the relative contribution of shared prompt difficulty versus shared answer format cannot be separated from this analysis.
Observational empirical analysis of completion groups. Completions from Qwen2.5-1.5B model evaluated on MATH, GSM8K, and DeepMath-103K datasets; grouped by prompt (eight completions per prompt).. n = 24,998.
Pooled within-group verifier-error correlation of 0.530 (95% CI: 0.500–0.560) across 24,998 groups from Qwen2.5-1.5B on MATH, GSM8K, and DeepMath-103K Design-effect-adjusted effective sample size of 1.70 for eight-completion groups under an exchangeable-error model Substantial variation in dependence by answer form: fractions, radicals, symbolic expressions, and intervals exhibit stronger clustering than unit annotations and percent signs
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
An empirical analysis of verifier error dependence in a large dataset from a single model, identifying a real statistical phenomenon (within-group correlation of 0.530) that challenges an independence assumption in RLVR; the work is descriptive and exploratory rather than hypothesis-testing, has not been peer reviewed, and does not evaluate a clinical or therapeutic intervention.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Group-based reinforcement learning with verifiable rewards (RLVR) scoresmultiple completions per prompt using automatic verifiers. Analysesbased on independent verifier errors may overlook dependence associatedwith shared answer formats. We investigate this dependence in24,998 groups of eight completions generated by Qwen2.5-1.5B onMATH, GSM8K, and DeepMath-103K. We estimate a pooled within-groupverifier-error correlation of 0.530 (95% confidence interval:0.500--0.560). Under an exchangeable-error model, this correspondsto a design-effect-adjusted effective sample size of 1.70 for aneight-completion group. Dependence varies substantially across answerforms: fractions, radicals, symbolic expressions, and intervals exhibitstronger clustering than unit annotations and percent signs. Replayinggroup-relative advantages across four rule-based verifier configurationsidentifies at least one advantage-sign disagreement in up to 0.83% ofgroups. Because a group is repeated sampling for one prompt, thiswithin-group clustering may reflect shared prompt difficulty as well asshared answer form, and we do not attempt to separate the two here.Unlike studies of correlated judgments across multiple evaluators, ouranalysis examines dependence across completions scored by the sameverifier. These findings motivate prompt- and answer-form-aware analysesof verifier noise rather than characterizations based solely onaggregate error rates.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.