Life sciences · Preprint
arXiv · September 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint reports that precision in detector-defined datasets is governed by true-positive prevalence in source pools via Bayes' theorem, not detector quality alone. Across three pools from a single detector, phantom contamination rates ranged from 0.0% to 81.7%, causing a +422% prediction error when precision transfer was attempted without accounting for prevalence. The contamination acts as a second structured signal rather than random noise, and can inflate or dilute estimators depending on whether the denominator is also contaminated.
Observational analysis with independent validation cohort. Events detected by a single instrument and independently validated as real or phantom; three source pools with contamination rates ranging from 0% to 81.7%. n = 781.
Phantom contamination rates across three pools: 81.7%, 9.0%, and 0.0% Precision transfer prediction error: +422% (predicted 0.955 vs measured 0.183) Bayes expression predicted all three phantom rates within 3.3%
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
Researchers and practitioners building machine-learning datasets via detectors should not assume detector precision transfers across source populations with different event prevalence. Contamination can systematically bias estimators in opposite directions depending on estimator design; Bayes' theorem can predict contamination rates if prevalence is known.
Single-instrument observational study demonstrating a methodological problem in detector-defined datasets with real data, but lacks the experimental validation, generalizability testing, or peer review needed to guide practice.
As stated by the source record.
Quoted from the source exactly as published.
Researchers and practitioners building machine-learning datasets via detectors should not assume detector precision transfers across source populations with different event prevalence. Contamination can systematically bias estimators in opposite directions depending on estimator design; Bayes' theorem can predict contamination rates if prevalence is known.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Many ML datasets are constructed by running a detector, heuristic, or model over candidate pools; accepted items become labels. Dataset precision is then governed by true-positive prevalence in each pool via Bayes, not solely by detector quality. Using one instrument and period, we hold a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom. One detector, three pools yield phantom rates 81.7%, 9.0%, and 0.0%. Transferring precision from the two high-rate pools to the low-rate pool predicts 0.955 versus measured 0.183, a +422% error; the Bayes expression predicts all three within 3.3%. The detected response curve is an exact convex combination of a true-event and a phantom component (residual 1.1e-16), with phantoms outnumbering true events 473 to 308, so contamination is a second signal with detector-inherited shape, not additive noise. Contamination direction depends on the estimator: on identical windows one statistic is diluted and another inflated because its denominator is also contaminated. A common normalization turns the estimator into a mean of ratios whose expectation need not exist; on the same 335 events it returns 0.40 where the well-defined estimator returns 0.10.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.