परिचय
जुलाई 2026 का benchmark, जो जाँचता है कि multimodal language models visual designs से executable multi-view visualization interfaces बना सकते हैं या नहीं। MV-Bench Tableau workbooks को structured ground truth की तरह इस्तेमाल करता है और visual fidelity, data bindings, coordinated interactions तथा code executability का मूल्यांकन करता है।
सारांश
1. Single-screen appearance से आगे: कई linked views वाले coordinated dashboards को test करता है
2. Executable benchmark: Code, rendered interfaces, datasets और interaction annotations के साथ 1,048 verified instances देता है
3. Structured ground truth: 92 Tableau workbooks से bindings और interactions निकालता है
4. Semantic evaluation: Visual fidelity के साथ data correctness और interaction completeness भी मापता है
5. महत्वपूर्ण capability gap: सबसे मजबूत tested model 75.45% layout accuracy तक पहुँचता है, लेकिन data-binding में केवल 21.71% और interaction accuracy में 11.68% हासिल करता है