Life sciences · Preprint
arXiv · August 19, 2026
Early or partial results. Treat as a signal, not a conclusion.
SparsePR is a training-free sparse attention method that reconstructs residual attention for video transformers, achieving 1.48x–2.61x end-to-end speedups while maintaining generation quality at 22.0–26.0% executed-pair density across four tested models. The work is a technical contribution to efficient transformer inference but has not undergone peer review and does not address clinical utility or generalization beyond the tested architectures.
Empirical algorithm evaluation with ablation study across four heterogeneous models. Four video generation and world-model transformer architectures; no human subjects or clinical population.. Intervention: SparsePR: training-free sparse attention combining Response-Coupled Partitioning and Probe-Fitted Residual Reconstruction.. Compared with: Ablated variants (without probe fitting, without response-coupled partitioning; implicit comparison to dense attention)..
SparsePR achieves 1.48x–2.61x end-to-end speedups across four heterogeneous video generation and world models. Generation quality preserved at 22.0–26.0% realized executed-pair density (i.e., 74–78% sparsity). Probe fitting accounts for most attention-reconstruction error reduction in ablation studies.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a methods/algorithm paper demonstrating a training-free optimization technique for video transformers with measured speedups and quality preservation, but lacks clinical or patient-outcome validation and has not been peer reviewed.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. Queries sharing a block route may have poorly overlapping supports, while retained attention mass alone does not determine the post-softmax error from skipped interactions. We show that partition geometry affects both pooled support and the predictability of the remaining residual from the sparse output. We introduce SparsePR, which combines Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction. Sampled-query key responses form paired K/V groups, whose centroids induce query-response coordinates for shared routing. A small set of exact query rows then calibrates a call-specific affine correction from the sparse output within the output subspace observed in the probe residuals. Across four heterogeneous video generation and world models, SparsePR consistently reduces attention-reconstruction error. Ablations show that probe fitting accounts for most of this reduction, while response-coupled partitioning lowers hard-drop error and improves reconstruction under a finite probe budget. SparsePR preserves generation quality at 22.0-26.0% realized executed-pair density while achieving 1.48x-2.61x end-to-end speedups. Project page: https://pardistaghavi.github.io/SparsePR-website/
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.