Life sciences · Preprint
arXiv · August 13, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint introduces SciFigBench, a diagnostic benchmark with 250 annotated scientific figures and over 34,000 evaluation setups designed to assess vision-language models on perception, reasoning, and behavioral reliability under uncertainty. The study reveals that models with comparable high reasoning accuracy (e.g., GPT-5.2 at 78.4% vs. Gemini 3.1 Pro at 81.0%) exhibit markedly different behaviors when facing missing or misleading visual evidence, suggesting that accuracy metrics alone do not predict reliability in deployments where uncertainty acknowledgment is critical.
Diagnostic benchmark evaluation study (preprint). Vision-language models (GPT-5.2, Gemini 3.1 Pro, and others) tested on 250 scientific figures with systematic variations and stress tests.. Intervention: Systematic presentation of scientific figures (original and transformed) with reasoning questions, resistance probes, caption-bias probes, and selective-blur manipulations to stress-test model behavior.. Compared with: Multiple VLMs compared on description quality (MQM), reasoning accuracy, hallucination rate, uncertainty admission, and resistance score..
GPT-5.2 achieves highest description quality (MQM 91.6) and reasoning accuracy (78.4%) but hallucinates unreadable content in 96% of cases. Gemini 3.1 Pro achieves comparable description quality (MQM 90.2) and higher reasoning accuracy (81.0%), but admits uncertainty in 71% of cases with strongest resistance score (0.91). SciFigBench contains 250 figures with 600+ hours of human annotation, extended to 34,000+ evaluation setups through image transformations, reasoning questions, and resistance probes.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This work is not directly applicable to clinical practice. However, for life-science professionals deploying VLMs in research or diagnostic workflows, the findings highlight that high accuracy on standard benchmarks does not guarantee safe behavior under uncertainty—a gap that must be evaluated before clinical or high-stakes scientific adoption.
A diagnostic benchmark study introducing a new evaluation framework for VLM behavior under uncertainty; reports descriptive findings and model comparisons without randomization, control, or clinical outcome validation, appropriate for a preprint stress-testing methodology.
As stated by the source record.
Quoted from the source exactly as published.
This work is not directly applicable to clinical practice. However, for life-science professionals deploying VLMs in research or diagnostic workflows, the findings highlight that high accuracy on standard benchmarks does not guarantee safe behavior under uncertainty—a gap that must be evaluated before clinical or high-stakes scientific adoption.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading). We introduce SciFigBench, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty. It contains 250 figures with high-quality human annotations across three evaluation aspects, totaling 600+ hours of annotation effort. We further extend these figures via image transformations, reasoning questions, resistance probes, caption-bias probes, and confirmed selective-blur targets, producing over 34,000 evaluation setups for stress testing. We further propose the Admittance-Resistance-Inductance (A-R-I) framework to evaluate whether models acknowledge insufficient evidence, resist misleading context, and infer cautiously from partial information. Our results reveal substantial behavioral differences among models. GPT-5.2 achieves the highest description quality (MQM 91.6) with strong reasoning accuracy (78.4%), yet hallucinates unreadable content in 96% of cases, whereas Gemini 3.1 Pro, a comparably capable model (MQM 90.2, reasoning 81.0%), admits uncertainty in 71% of such cases and achieves the strongest resistance score (0.91). These findings show that high perception and reasoning accuracy alone do not guarantee behavioral reliability, a dimension critical for deployment in scientific workflows.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.