Life sciences · Preprint
arXiv · September 10, 2026
Raises a question worth testing. It does not answer one.
This unreviewed study reports that language models exhibit markedly different susceptibility to output collapse under recursive training—ranging from 0.187 to 0.940 unique 4-grams after five generations—and that this fragility is an intrinsic checkpoint property predictable by brief self-iteration, independent of parameter scale. The work characterizes a phenomenon but does not establish causation, validate prediction utility in deployment, or measure real-world consequences.
Controlled computational ecosystem study with recursive contamination protocol. 13 publicly released language-model checkpoints of unspecified architectures and training origins.. Intervention: Recursive contamination (model-generated text fed back into training corpus); top-p tightening at generation time; data-side filtering.. Compared with: Baseline recursive contamination protocol without intervention; human-only text baseline (implicit).. n = 13.
Unique 4-gram outcome after five generations ranges from 0.187 to 0.940 across 13 checkpoints, a roughly five-fold spread. Spearman correlation of model ordering remains stable at 0.91–0.97 when corpus composition changes and 0.93–0.98 across random seed variation. Parameter scale alone does not explain fragility; a three-size ladder within one model family is non-monotonic in size.
Diversity metric (4-gram uniqueness) is a surrogate; no measurement of downstream task performance, utility, or harm.
The source did not state who this applies to in practice.
An exploratory computational study characterizing model-specific fragility under recursive training using descriptive metrics and correlational analysis, without causal intervention or clinical outcomes; raises questions about checkpoint properties rather than answering them definitively.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Model-generated text is finding its way back into training corpora, and there is plenty of evidence that training on such data over and over collapses output diversity. Prior work has studied the phenomenon itself: which protocols and which data mixtures cause collapse. But different models behave very differently under the same process. We fix one recursive contamination protocol and let 13 publicly released checkpoints form an ecosystem that shares a common corpus for five generations. The unique 4-gram outcome after five generations ranges from 0.187 to 0.940 across checkpoints, a roughly five-fold spread: some models are barely touched, others degenerate into repetitive fragments. Changing the composition of the shared pool or mixing in human text keeps the Spearman correlation of the ordering at 0.91--0.97, and changing the random seed keeps it at 0.93--0.98. Whether a model collapses easily under recursive training is, then, a property of the checkpoint itself, and one that has gone largely unexamined. Parameter scale alone does not explain it, since a three-size ladder within one family is not monotonic in size, and none of the static indicators we tested predicts it either. What does work is cheap: let a model iterate on its own output for two or three generations, and its fragility in the larger ecosystem can be inferred from that alone. Collapse speed also responds to intervention. Tightening top-p, which cuts the low-probability tail at generation time, nearly stops collapse within three generations and stabilizes six checkpoints spanning the whole spectrum together, while data-side filtering slows collapse without stopping it.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.