List पर वापस जाएँ
📄Papers

MV-Bench: Benchmarking MLLMs for Coordinated Multi-View Interface Construction

Zhao Y. et al.
2026-07

परिचय

जुलाई 2026 का benchmark, जो जाँचता है कि multimodal language models visual designs से executable multi-view visualization interfaces बना सकते हैं या नहीं। MV-Bench Tableau workbooks को structured ground truth की तरह इस्तेमाल करता है और visual fidelity, data bindings, coordinated interactions तथा code executability का मूल्यांकन करता है।

सारांश

MV-Bench क्या नया संकेत देता है:

1. Single-screen appearance से आगे: कई linked views वाले coordinated dashboards को test करता है
2. Executable benchmark: Code, rendered interfaces, datasets और interaction annotations के साथ 1,048 verified instances देता है
3. Structured ground truth: 92 Tableau workbooks से bindings और interactions निकालता है
4. Semantic evaluation: Visual fidelity के साथ data correctness और interaction completeness भी मापता है
5. महत्वपूर्ण capability gap: सबसे मजबूत tested model 75.45% layout accuracy तक पहुँचता है, लेकिन data-binding में केवल 21.71% और interaction accuracy में 11.68% हासिल करता है

Tags

ui-generationmultimodal-modelsdata-visualizationinteractionbenchmark
मूल लेख पढ़ें

मिलते-जुलते लेख

📄Papers

LEGOUI: Designing with UI-DSL Bricks for Transparent, Controllable Generation

अगस्त 2026 का HCI paper, जो शुरुआती design ideation के लिए staged Generative UI framework पेश करता है। LEGOUI prompt से मिले और model द्वारा अनुमानित फैसलों को provenance-aware UI DSL में दर्ज करता है, users को layout, interaction, relation और style stages में फैसले स्वीकार, अस्वीकार या जोड़ने देता है, और बदलती specification से interface render करता है।

genuiui-dsldesign-time
लेखक Zhou Y. et al. · 2026-08
और पढ़ें
📄Papers

Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation

जुलाई 2026 का paper, जो ESPP नाम की evaluation method पेश करता है। यह एक LLM judge की जगह evidence-grounded और मनोवैज्ञानिक रूप से विविध personas का panel रखता है। Panelists generated-interface screenshots को स्वतंत्र रूप से rate करते हैं, bounded-confidence mechanism के जरिए राय साझा करते हैं, और उनके scores को social weighting के साथ जोड़ा जाता है।

genuiui-evaluationhuman-preferences
लेखक Wu Z. et al. · 2026-07
और पढ़ें
📄Papers

Design Theater: A Benchmark for Generative UI

जुलाई 2026 का paper, जो यह जाँचने के लिए benchmark पेश करता है कि Generative UI tools users को बताए गए design rationales सच में लागू करते हैं या नहीं। पाँच tools से बने 120 interfaces और 24 structural, styling व functional tasks में यह study बताए गए design decisions और generated result के बीच अंतर मापती है।

genuiui-evaluationdesign-rationale
लेखक Imteyaz K. et al. · 2026-07
और पढ़ें

Generative UI resources का curated संग्रह।

बनाया गया ❤️ community द्वारा