Life sciences · Preprint
arXiv · September 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
LayerWiseBench is a new structured benchmark for evaluating visual language models and image editors on layer-wise chart understanding and editing tasks. Qwen3.5-27B achieves high accuracy on layer attribution (93.04%) and binding (97.46%) but struggles with visibility ordering (61.46%), while all tested image editors show poor performance on visibility-constrained edits (mIoU 0.37–2.00%), identifying front-to-back component relations as a systematic weakness.
Benchmark construction and descriptive model evaluation. Synthetic charts across 14 chart paradigms generated from executable chart programs; no human subjects.. Intervention: LayerWiseBench evaluation tasks (layer attribution, binding, visibility ordering, and editing). n = 2,800.
Qwen3.5-27B achieves 93.04% accuracy on layer attribution and 97.46% on layer binding Qwen3.5-27B achieves only 61.46% accuracy on visibility ordering Four evaluated image editors show overall mIoU ranging from 1.49% to 4.93%
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a benchmark paper introducing a new evaluation dataset and reporting descriptive performance metrics from model evaluations, without comparative clinical trials, interventional studies, or peer-reviewed publication.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Charts are structured visual compositions whose elements have distinct functional roles, semantic correspondences, and visibility relations. This structural view motivates evaluating whether models can understand and manipulate charts at the layer level. Existing chart benchmarks, however, primarily assess the correctness or fidelity of final outputs and do not directly evaluate these layer-wise behaviors. We present LayerWiseBench, a benchmark organized around three core concepts, layer attribution, layer binding, and visibility ordering, that structure its chart-understanding and chart-editing evaluations. Generated from executable chart programs, LayerWiseBench pairs each rendered chart with spatially aligned per-layer RGBA assets and construction-derived labels for functional roles, semantic bindings, and visibility relations. From this layer-wise representation, we derive controlled understanding questions, editing targets, reference images, and evaluation regions. It contains 2,800 source charts across 14 chart paradigms, from which we derive 7,329 layer-wise understanding questions and 53,791 instruction-guided editing variants. Among the evaluated VLMs, Qwen3.5-27B, which achieves the highest QA macro-average, obtains 93.04% accuracy on layer attribution and 97.46% on layer binding, but only 61.46% on visibility ordering. Across the four evaluated image editors, overall mIoU ranges from 1.49% to 4.93%, and visibility-constrained edits have the lowest mIoU for every editor, ranging from 0.37% to 2.00%. Taken together, these results identify tasks involving front-to-back relations between overlapping components as a recurring challenge across understanding and editing, motivating more explicit modeling of component identity and visibility relations.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.