Life sciences · Preprint
arXiv · September 5, 2026
Early or partial results. Treat as a signal, not a conclusion.
DataFlex-RL is an unreviewed evaluation platform for comparing data policies in reinforcement learning with verifiable rewards. Across 13 configurations tested on two models and 12 benchmarks, no rollout-selection, reweighting, or adaptive mixture method achieved a statistically significant improvement over uniform GRPO, and findings did not replicate consistently between models. The work demonstrates that data policies affect training but does not identify a reproducibly superior approach in the conditions studied.
Controlled comparative evaluation study with matched seeds across multiple models and benchmark suites. Two pretrained language models (Qwen2.5-7B-Base and Llama-3.1-8B-Base) trained on 12 mathematics, logic, and science benchmarks using RLVR. An additional nine Qwen2.5-7B-Instruct runs were rescored for evaluation sensitivity analysis.. Intervention: Thirteen data policy configurations comprising eight rollout-selection or reweighting methods and three adaptive mixture methods for selecting and weighting training rollouts under GRPO.. Compared with: Uniform GRPO (baseline) and fixed equal mixture across domains..
Uniform GRPO improves domain-balanced average accuracy by 7.76 percentage points over untrained checkpoint None of eight rollout-selection or reweighting methods achieves paired 95% CI excluding zero relative to uniform sampling None of three adaptive mixtures outperforms fixed equal mixture at 95% CI precision level
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unreviewed computational study comparing data policy configurations in a controlled setting with null findings; it advances methodology for RLVR evaluation but does not demonstrate that any proposed approach outperforms the baseline.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary experiment evaluates 13 configurations across 12 matched seeds using Qwen2.5-7B-Base and 12 mathematics, logic, and science benchmarks. Uniform GRPO improves the domain-balanced average accuracy by 7.76 percentage points over the untrained checkpoint. None of the eight rollout-selection or reweighting methods achieves a paired 95% confidence interval that excludes zero relative to uniform sampling, and none of the three adaptive mixtures outperforms a fixed equal mixture at the same level of precision. A corrected 12-seed extension on Llama-3.1-8B-Base places the additional methods on the same score scale as the original controls, but does not reveal a consistent winner in terms of observed mean performance. We also quantify evaluation sensitivity by rescoring nine Qwen2.5-7B-Instruct runs using a math-heavy six-benchmark summary, consisting of five mathematics benchmarks and GPQA-Diamond but no logic benchmark, and comparing it with the domain-balanced 12-benchmark summary. The resulting rankings are negatively correlated, with a correlation coefficient of -0.33, whereas summaries that retain all 12 benchmarks largely agree. Across the controlled settings studied here, changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.