AUG 13, 2026 · PREPRINT
How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
arXiv
A diagnostic benchmark study introducing a new evaluation framework for VLM behavior under uncertainty; reports descriptive findings and model comparisons without randomization, control, or clinical outcome validation, appropriate for a preprint stress-testing methodology.
Reported
GPT-5.2 description quality (MQM)91.6
GPT-5.2 reasoning accuracy78.4%
GPT-5.2 hallucination rate on unr…96%