العودة إلى القائمة
📄الأبحاث

EvoGenUI-Bench: Evaluating Multi-Turn Generative UI Assistants

Peng Y. et al.
2026-08

حول

August 2026 benchmark testing whether LLM assistants can maintain one executable web interface as user requirements evolve. EvoGenUI-Bench contains 150 five-turn tasks across information presentation, stateful interaction, and tool-grounded external state, with browser-based evaluation over screenshots, DOM and source evidence, actor traces, and runtime logs.

الملخص

Why EvoGenUI-Bench Adds New Signal:

1. Evaluates interface evolution: Tests cumulative revisions to one artifact instead of isolated page generation
2. Measures requirement retention: Introduces Adjacent Pass Retention to show whether a correct interface stays correct after the next request
3. Executes every artifact: Grounds scoring in browser behavior, rendered output, source, DOM, traces, logs, and external-state readback
4. Exposes a reliability gap: The strongest model reaches 74.9% turn success but completes only 37.3% of five-turn episodes
5. Diagnoses concrete failures: Separates information-architecture, derived-state, affordance-binding, external-grounding, and requirement-decomposition breakdowns

الوسوم

generative-uimulti-turnbenchmarkstate-managementbrowser-evaluation

مقالات ذات صلة

📄الأبحاث

Toward Frontier-Quality Declarative UI Generation at Small-Model Cost

September 2026 Harness4GenUI paper studying whether small language models can generate production-oriented A2UI from approved component catalogs. Across task-management and cloud-console domains, the authors compare training-data strategies, model sizes, catalog sizes, cost, latency, binding correctness, and rendered quality.

a2uismall-language-modelssupervised-fine-tuning
بواسطة Yang Y. et al. · 2026-09
اقرأ المزيد
📄الأبحاث

Maru: Information Architecture for Aligned, Persistent Generative UI

August 2026 HCI paper introducing information architecture as shared state for Generative UI. Maru captures how users partition, order, name, and prioritize information, turns those choices into persistent rules, and applies them across successive interface generations instead of rebuilding structure from each prompt.

generative-uiinformation-architecturepersonalization
بواسطة Kim E. et al. · 2026-08
اقرأ المزيد
📄الأبحاث

SchemaGUI: A Schema-Driven Benchmark for Controllable GUI Generation

August 2026 benchmark for evaluating LLM-generated interfaces with deterministic annotations instead of noisy human labels. SchemaGUI synthesizes bilingual instructions and function-call references from parameterized UI schemas, then measures feasibility, geometry, and layout complexity across six scenarios and five models.

ui-generationbenchmarkschema-driven
بواسطة Dong J. et al. · 2026-08
اقرأ المزيد

مجموعة منتقاة من مصادر Generative UI وGenUI وواجهات AI الديناميكية.

صُنع باستخدام ❤️ من قِبل المجتمع