List पर वापस जाएँ
📄Papers

Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation

Wu Z. et al.
2026-07

परिचय

जुलाई 2026 का paper, जो ESPP नाम की evaluation method पेश करता है। यह एक LLM judge की जगह evidence-grounded और मनोवैज्ञानिक रूप से विविध personas का panel रखता है। Panelists generated-interface screenshots को स्वतंत्र रूप से rate करते हैं, bounded-confidence mechanism के जरिए राय साझा करते हैं, और उनके scores को social weighting के साथ जोड़ा जाता है।

सारांश

ESPP क्या नया संकेत देता है:

1. बहु-दृष्टिकोण evaluation: फैसले को एक छिपे हुए viewpoint में समेटने के बजाय दिखाता है कि अलग user groups एक ही generated interface को कैसे देखते हैं
2. Human-alignment evidence: Human ratings के साथ Pearson correlation को 0.716 से 0.922 तक सुधारता है
3. Controlled comparison: दिखाता है कि prompt ensemble इस सुधार का केवल लगभग एक-तिहाई ही हासिल करता है
4. Dimension-level disagreement: Personas की ranking पर सहमति के साथ उन खास interface qualities को भी बचाए रखता है जिन पर उनकी राय अलग है
5. Reproducible implementation: चलने योग्य three-stage scoring pipeline, sample screenshots और worked cases जारी करता है

Tags

genuiui-evaluationhuman-preferencesllm-as-a-judgepersonas
मूल लेख पढ़ें

मिलते-जुलते लेख

📄Papers

LEGOUI: Designing with UI-DSL Bricks for Transparent, Controllable Generation

अगस्त 2026 का HCI paper, जो शुरुआती design ideation के लिए staged Generative UI framework पेश करता है। LEGOUI prompt से मिले और model द्वारा अनुमानित फैसलों को provenance-aware UI DSL में दर्ज करता है, users को layout, interaction, relation और style stages में फैसले स्वीकार, अस्वीकार या जोड़ने देता है, और बदलती specification से interface render करता है।

genuiui-dsldesign-time
लेखक Zhou Y. et al. · 2026-08
और पढ़ें
📄Papers

Design Theater: A Benchmark for Generative UI

जुलाई 2026 का paper, जो यह जाँचने के लिए benchmark पेश करता है कि Generative UI tools users को बताए गए design rationales सच में लागू करते हैं या नहीं। पाँच tools से बने 120 interfaces और 24 structural, styling व functional tasks में यह study बताए गए design decisions और generated result के बीच अंतर मापती है।

genuiui-evaluationdesign-rationale
लेखक Imteyaz K. et al. · 2026-07
और पढ़ें
📄Papers

MV-Bench: Benchmarking MLLMs for Coordinated Multi-View Interface Construction

जुलाई 2026 का benchmark, जो जाँचता है कि multimodal language models visual designs से executable multi-view visualization interfaces बना सकते हैं या नहीं। MV-Bench Tableau workbooks को structured ground truth की तरह इस्तेमाल करता है और visual fidelity, data bindings, coordinated interactions तथा code executability का मूल्यांकन करता है।

ui-generationmultimodal-modelsdata-visualization
लेखक Zhao Y. et al. · 2026-07
और पढ़ें

Generative UI resources का curated संग्रह।

बनाया गया ❤️ community द्वारा