Life sciences · Preprint
arXiv · September 8, 2026
Posted before peer review. The findings may change or fail to hold.
ActionSplice is a proposed inference framework for editing actions in chunk-autoregressive video world models without replaying computations. It reports relative reductions in LPIPS (56–76%) and speedup multiples (1.69–2.73×) on synthetic benchmarks across two models, but these are computational metrics on unevaluated video generation tasks without peer review or external validation.
Preprint. Intervention: ActionSplice inference framework using Counterfactual State Transport (CST) with retargeting (CST_R) and temporal-splicing (CST_T) variants. Compared with: Direct condition swapping and waiting baseline.
CST_R reduces rollback-relative LPIPS by 61.5% and 75.9% relative to direct condition swapping on minWM-Wan Action2V and HY-WM1.5 respectively CST_T reduces suffix LPIPS by 56.1% and 77.5% while providing 2.73× and 1.69× pixel-ready speedups over waiting Under HY-WorldPlay protocol, CST_R obtains PSNR of 25.66 dB, SSIM of 0.6902, and LPIPS of 0.1337
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unrefereed arXiv preprint describing a technical framework for video world models; it reports computational benchmarks on synthetic video generation tasks but lacks peer review and clinical or real-world validation.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Chunk-autoregressive video world models typically condition each generated chunk on one action. An action received during sampling must therefore wait for the next chunk, condition future solver evaluations on a state produced under the previous action, or trigger rollback that repeats completed evaluations. We introduce ActionSplice, an inference framework that formulates this problem as Counterfactual State Transport (CST). A lightweight corrector transports the interrupted backbone-native representation toward the matched state induced by the revised action at the same solver step. The world model and sampler remain frozen, and sampling resumes without replaying completed evaluations. The retargeting variant $\mathrm{CST}*{R}$ updates the entire active chunk, while the temporal-splicing variant $\mathrm{CST}*{T}$ preserves a temporal prefix and updates only the suffix. Across minWM-Wan Action2V and HY-WM1.5, $\mathrm{CST}*{R}$ reduces rollback-relative LPIPS by 61.5% and 75.9% relative to direct condition swapping. $\mathrm{CST}*{T}$ reduces suffix LPIPS by 56.1% and 77.5%, respectively, while providing $2.73\times$ and $1.69\times$ pixel-ready speedups over waiting. Under the HY-WorldPlay protocol, $\mathrm{CST}_{R}$ obtains a PSNR of 25.66 dB, an SSIM of 0.6902, and an LPIPS of 0.1337 against the original rollout.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.