Life sciences · Preprint
arXiv · August 13, 2026
Posted before peer review. The findings may change or fail to hold.
This preprint introduces TraVEL, a trajectory-guided fine-tuning method for adapting multimodal video embeddings to driving-video retrieval tasks. The method reports improvements in motion-centric retrieval metrics (longitudinal and lateral mAP gains) across two model scales, but lacks peer review and does not address clinical or real-world safety validation.
Methods development and benchmark evaluation. Driving video clips from the nuReasoning dataset with paired reasoning traces and trajectory annotations.. Intervention: TraVEL: trajectory-guided fine-tuning framework using ego-trajectory similarity as reward within Group Relative Policy Optimization.. Compared with: Supervised fine-tuning (SFT) on captions alone using InfoNCE objective..
TraVEL raises longitudinal mAP by 9.8 points at 2B model scale and 7.2 points at 8B model scale relative to supervised fine-tuning (SFT) baseline TraVEL raises lateral mAP by 4.7 points at 2B and 1.5 points at 8B relative to SFT Motion-aware fine-tuning using ego-trajectory similarity as reward outperforms caption supervision alone for fine-grained motion understanding
Applicability to safety-critical driving systems or real-world deployment is not evaluated or discussed.
The source did not state who this applies to in practice.
This is an unrefereed arXiv preprint describing a machine learning method for video retrieval; it reports technical performance metrics but has not undergone peer review.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule-based retrieval systems can explicitly target driving events, but typically require expert-defined rules, auxiliary data, and multi-stage perception pipelines. Multimodal embedding models offer a simpler and more efficient alternative by representing each video with a single searchable vector. However, general-purpose models often rely on shortcuts from static scene context and struggle to distinguish motion-centric events, such as turning left versus right or accelerating versus decelerating. In this work, we study how to adapt a general-purpose multimodal embedding model to driving-video retrieval. We first fine-tune Qwen3-VL-Embedding on paired clips and reasoning traces from nuReasoning using an InfoNCE objective. While this stage substantially improves overall retrieval, caption supervision alone remains insufficient for fine-grained motion understanding. We therefore introduce TraVEL (Trajectory-Guided Video Embedding Learning), a motion-aware fine-tuning framework that uses ego-trajectory similarity as a reward within Group Relative Policy Optimization. Trajectories serve only as privileged training supervision; retrieval still operates on single-vector video embeddings without ego poses, expert rules, or auxiliary perception outputs. We further construct a driving-video retrieval benchmark from nuReasoning. Experiments show that TraVEL improves motion-centric retrieval across model scales: relative to SFT, it raises longitudinal and lateral mAP by 9.8 and 4.7 points at 2B, with corresponding gains of 7.2 and 1.5 points at 8B. TraVEL thus combines physically grounded supervision with efficient embedding-based search.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.