Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
ALIGN-HOLD is a machine-learning system for real-time matching control in ride-hailing, validated via a 28-day randomized A/B experiment on DiDi's platform. The intervention showed statistically significant improvements in trip completion rate, driver income, and passenger cancellations compared to the existing production policy, but the preprint does not report detailed effect sizes, confidence intervals, or p-values.
Randomized A/B experiment in production. Approximately 100,000 passenger requests per day on DiDi's ride-hailing platform.. Intervention: ALIGN-HOLD: experience alignment framework that learns hold policy from implicit marketplace preferences using a Reward Model.. Compared with: Deployed production policy (EXHOLD), which uses bandit-based policies from handcrafted reward combinations.. DiDi's Brazil marketplace..
Statistically significant improvement in trip completion rate compared with deployed production policy Statistically significant increase in driver income compared with deployed production policy Statistically significant reduction in passenger cancellations before and after driver acceptance
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A production A/B experiment on a real ride-hailing platform with statistically significant improvements in key metrics, but the work is a preprint, lacks peer review, reports only summary results without detailed effect sizes or confidence intervals, and is described as an engineering system deployment rather than a clinical or definitive health/safety study.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Real-time hold control is a high-leverage mechanism in large-scale ride-hailing systems: by selectively deferring driver-order pairs, the platform can wait for better matching opportunities and improve end-to-end passenger-driver experience. Existing production systems such as EXHOLD learn bandit-based hold policies from handcrafted combinations of trip completion, cancellations, waiting time, and driver effort. However, designing such rewards becomes increasingly difficult as marketplace preferences are heterogeneous and observed passenger-driver behavior can be sparse, noisy, and affected by dynamic supply-demand conditions. We present ALIGN-HOLD, a production-scale experience alignment framework that learns hold policy from implicit marketplace preferences. ALIGN-HOLD constructs complementary preference pairs from order trajectories, driver trajectories, and contemporaneous local matching graphs, and trains an experience Reward Model (RM) using balanced multi-view sampling and model-adaptive hard preference sampling. During simulator-based policy learning, the frozen RM provides a dense, context-dependent reward and supports label-free filtering of low-identifiability interactions whose behavioral feedback is difficult to attribute to matching quality. We deploy ALIGN-HOLD on DiDi's ride-hailing platform and evaluate it in a 28-day randomized A/B experiment, covering approximately 100,000 passenger requests per day. Compared with the deployed production policy, ALIGN-HOLD achieves statistically significant improvements in trip completion rate and driver income, while significantly reducing passenger cancellations before and after driver acceptance. Complementary ablations, RM diagnostics, and behavioral analyses validate the contributions of the proposed components. ALIGN-HOLD has been fully ramped up and is currently serving DiDi's Brazil marketplace.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.