Life sciences · Preprint
arXiv · August 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
MMDiff is a novel framework for training multimodal sparse autoencoders to identify and control interpretable features in multimodal large language models. Feature removal achieved selective task-specific degradation (12–17% on spatial and OCR tasks, 24% on safety attacks) and steering improved accuracy modestly (+1.8–3.6%), demonstrating proof-of-concept for interpretability and targeted control.
Single-arm comparative framework validation study. Three multimodal large language model families: LLaVA-MORE, PaliGemma 2, and InternVL3.5.. Intervention: MMDiff multimodal model-diffing framework with feature removal and steering applied to identified feature directions.. Compared with: Single-layer steering baseline for steering experiments; no comparator for removal experiments..
Feature removal selectively degraded target behaviors by average of 12% on spatial tasks and 17% on OCR Feature removal reduced multimodal safety attack success rate by 24% Feature steering improved spatial and OCR accuracy by +3.6% and +1.8% on average over single-layer steering baseline
Feature removal reduced multimodal safety attack success rate by 24%
The source did not state who this applies to in practice.
This is an early-stage methodological study demonstrating proof-of-concept for a novel interpretability and control framework on three MLLM families, with modest effect sizes and no peer-review validation.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.