العودة إلى القائمة
📄الأبحاث

Design Theater: A Benchmark for Generative UI

Imteyaz K. et al.
2026-07

حول

July 2026 paper introducing a benchmark for checking whether Generative UI tools implement the design rationales they present to users. Across 120 interfaces from five tools and 24 structural, styling, and functional tasks, the study measures gaps between claimed design decisions and the generated result.

الملخص

Why Design Theater Adds New Signal:

1. Rationale-to-interface evaluation: Tests whether generated design explanations are visible in the actual UI
2. Three requirement classes: Separates structural, styling, and functional implementation quality
3. Cross-tool benchmark: Evaluates 120 interfaces produced by five Generative UI tools
4. Concrete failure evidence: Finds more than 25% of stated rationales unimplemented on average, rising to 34% for functional requirements
5. UX principle coverage: Measures whether tools recognize and implement the design principles embedded in prompts

الوسوم

genuiui-evaluationdesign-rationaleux-principlesbenchmark

مقالات ذات صلة

📄الأبحاث

LEGOUI: Designing with UI-DSL Bricks for Transparent, Controllable Generation

August 2026 HCI paper introducing a staged Generative UI framework for early design ideation. LEGOUI records prompt-derived and model-inferred decisions in a provenance-aware UI DSL, lets users accept, reject, or add decisions across layout, interaction, relation, and style stages, and renders from the evolving specification.

genuiui-dsldesign-time
بواسطة Zhou Y. et al. · 2026-08
اقرأ المزيد
📄الأبحاث

Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation

July 2026 paper introducing ESPP, an evaluation method that replaces one LLM judge with a panel of evidence-grounded, psychologically diverse personas. Panelists independently rate generated-interface screenshots, exchange opinions through a bounded-confidence mechanism, and are combined with social weighting.

genuiui-evaluationhuman-preferences
بواسطة Wu Z. et al. · 2026-07
اقرأ المزيد
📄الأبحاث

MV-Bench: Benchmarking MLLMs for Coordinated Multi-View Interface Construction

July 2026 benchmark for testing whether multimodal language models can generate executable multi-view visualization interfaces from visual designs. MV-Bench uses Tableau workbooks as structured ground truth and evaluates visual fidelity, data bindings, coordinated interactions, and code executability.

ui-generationmultimodal-modelsdata-visualization
بواسطة Zhao Y. et al. · 2026-07
اقرأ المزيد

مجموعة منتقاة من مصادر Generative UI وGenUI وواجهات AI الديناميكية.

صُنع باستخدام ❤️ من قِبل المجتمع