О материале
July 2026 benchmark for testing whether multimodal language models can generate executable multi-view visualization interfaces from visual designs. MV-Bench uses Tableau workbooks as structured ground truth and evaluates visual fidelity, data bindings, coordinated interactions, and code executability.
Краткое содержание
1. Beyond single-screen appearance: Tests coordinated dashboards with multiple linked views
2. Executable benchmark: Provides 1,048 verified instances with code, rendered interfaces, datasets, and interaction annotations
3. Structured ground truth: Derives bindings and interactions from 92 Tableau workbooks
4. Semantic evaluation: Measures data correctness and interaction completeness alongside visual fidelity
5. Important capability gap: The strongest tested model reaches 75.45% layout accuracy but only 21.71% data-binding and 11.68% interaction accuracy