Life sciences · Preprint
arXiv · August 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
This unreviewed methodological study demonstrates that standard rolling-origin evaluation protocols systematically misrepresent the atom structure (occurrence frequency) of time-series benchmarks, leading to reversed model rankings. The authors propose a matched protocol and show that an autoregressive hurdle model outperforms a conditional flow by up to 153× on five of six datasets, but model ordering remains inconsistent across different occurrence statistics.
Benchmarking study with controlled matched protocol over multiple seeds. Seven generative time-series models (including autoregressive hurdle and conditional flow) benchmarked on datasets with high point masses at a single value.. Intervention: Matched rolling-origin evaluation protocol that preserves dataset atom structure. Compared with: Standard rolling-origin protocol; five different occurrence statistics.
On one benchmark, dataset contains 42% zeros while evaluation windows contain 13%; on another 47% versus 5%, showing systematic protocol-data mismatch Rolling-origin evaluation protocol mismatch reversed conclusions about the strongest occurrence model in the study Autoregressive hurdle beats conditional flow on five of six datasets under matched protocol, by up to a factor of 153
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This work does not directly address clinical applications but is critical for practitioners relying on generative time-series models in any domain: standard evaluation protocols may systematically misrepresent model performance, potentially reversing true rankings. Users should adopt matched evaluation protocols that preserve dataset occurrence structure.
An unreviewed methodological study identifying critical evaluation problems in generative time-series models, proposing solutions but lacking peer review and presenting no definitive new model or algorithm.
As stated by the source record.
Quoted from the source exactly as published.
This work does not directly address clinical applications but is critical for practitioners relying on generative time-series models in any domain: standard evaluation protocols may systematically misrepresent model performance, potentially reversing true rankings. Users should adopt matched evaluation protocols that preserve dataset occurrence structure.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not rain, no ride is requested, no part is ordered. We report what happens when such data is evaluated carefully. First, the standard rolling-origin protocol can score a model on a window whose atom structure bears no resemblance to the dataset: on one benchmark the dataset is $42\%$ zeros and the evaluation windows are $13\%$, on another $47\%$ against $5\%$. This is not a cosmetic problem --- it reversed one of our own conclusions, turning the strongest occurrence model in our study into what looked like a cautionary tale. Second, we give a control in which CRPS is invariant \emph{by construction} while the temporal coupling is destroyed, which measures exactly how much that coupling contributes to a chosen statistic. Third, benchmarking seven models on a matched protocol over five seeds, an autoregressive hurdle beats a conditional flow on five of six datasets, by up to a factor of $153$, while the flow's own occurrence statistics vary by up to $62\%$ across training seeds and every baseline is deterministic. Finally, the model ordering is not the same under five different occurrence statistics, and the two that do not share a construction agree with each other least.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.