SEP 10, 2026 · PREPRINT
New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models
arXiv
This is a controlled benchmark study evaluating decision-making in vision language models on a novel task; it reveals limitations in reasoning but does not measure clinical or patient outcomes, and the findings are descriptive rather than hypothesis-testing.
Reported
Action repetition rate across six…95.1% to 100%
Best model correct joint decision…5.9%
Number of open models evaluated6