Life sciences · Preprint
arXiv · August 14, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint identifies forecast collapse—pathologically flat predictions and poor stock ranking—as a systematic failure of time-series foundation models on low-predictability targets like hourly equity returns. The authors attribute it to a calibration-ranking tradeoff and propose CalibRank, which reportedly triples cross-sectional correlation on their benchmark, but the work lacks peer review and independent validation.
Computational analysis; observational study across multiple models and benchmarks. 1,000 US equities with hourly returns; trading volume used as a contrasting low-collapse target.. Intervention: CalibRank objective function. Compared with: Conventional per-series squared-error minimization and direct cross-sectional correlation optimization.
Forecast collapse (poor stock ranking) observed when forecasting hourly returns for 1,000 US equities, but largely absent when forecasting trading volume. Phenomenon spans twelve deep-learning forecasting models and 97 public benchmark configurations, closely tied to target predictability. CalibRank objective nearly triples cross-sectional correlation on Finance1K while maintaining amplitude close to target.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
An unreviewed technical study identifying a phenomenon in foundation models with a proposed solution, lacking validation on held-out test data or independent benchmarks.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low predictability limits the amplitude of calibrated point forecasts, while per-series objectives leave cross-series structure unidentified. These findings reveal a calibration-ranking tradeoff: optimizing squared error leads to flat predictions, whereas directly optimizing cross-sectional correlation improves ranking but can inflate forecast amplitude by more than an order of magnitude. To address this tradeoff, we introduce CalibRank, a simple objective that balances calibration and ranking. On Finance1K, CalibRank nearly triples cross-sectional correlation while keeping amplitude close to the target, and improves correlation on all tested models. Our results reveal a blind spot in conventional time-series evaluation: per-series metrics can hide failures in cross-series structure needed by downstream decisions.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.