This is an early-stage methodological study demonstrating proof-of-concept for a novel interpretability and control framework on three MLLM families, with modest effect sizes and no peer-review validation.
Reported
Average degradation on spatial ta…12%
Average degradation on OCR (featu…17%
Attack success rate reduction (mu…24%