العودة إلى القائمة
📄الأبحاث

FlowEval: Reference-Based Evaluation of Generated User Interfaces

Wu J. et al.
2026-05

حول

Apple-affiliated research on evaluating whether generated interfaces support realistic user journeys. FlowEval runs validated tasks on reference websites and generated analogs, compares their navigation traces with reference-based similarity metrics, and tests how well those scores align with expert UI judgments.

الملخص

Why FlowEval Adds New Signal:

1. Interaction-first evaluation: Measures complete task flows instead of relying only on screenshots or code correctness
2. Trusted references: Compares generated sites with traces collected from known, high-quality interfaces
3. Automatable method: Uses computer-use agents and navigation tasks to collect comparable interaction sequences
4. Interpretable diagnostics: Trace similarity can reveal where a generated flow diverges from the reference experience
5. Human alignment evidence: Reports strong correlation between reference-based metrics and expert evaluations of generated UIs

الوسوم

ui-evaluationinteraction-flowscomputer-use-agentsbenchmarkhuman-evaluation

مقالات ذات صلة

📄الأبحاث

Toward Frontier-Quality Declarative UI Generation at Small-Model Cost

September 2026 Harness4GenUI paper studying whether small language models can generate production-oriented A2UI from approved component catalogs. Across task-management and cloud-console domains, the authors compare training-data strategies, model sizes, catalog sizes, cost, latency, binding correctness, and rendered quality.

a2uismall-language-modelssupervised-fine-tuning
بواسطة Yang Y. et al. · 2026-09
اقرأ المزيد
📄الأبحاث

EvoGenUI-Bench: Evaluating Multi-Turn Generative UI Assistants

August 2026 benchmark testing whether LLM assistants can maintain one executable web interface as user requirements evolve. EvoGenUI-Bench contains 150 five-turn tasks across information presentation, stateful interaction, and tool-grounded external state, with browser-based evaluation over screenshots, DOM and source evidence, actor traces, and runtime logs.

generative-uimulti-turnbenchmark
بواسطة Peng Y. et al. · 2026-08
اقرأ المزيد
📄الأبحاث

Maru: Information Architecture for Aligned, Persistent Generative UI

August 2026 HCI paper introducing information architecture as shared state for Generative UI. Maru captures how users partition, order, name, and prioritize information, turns those choices into persistent rules, and applies them across successive interface generations instead of rebuilding structure from each prompt.

generative-uiinformation-architecturepersonalization
بواسطة Kim E. et al. · 2026-08
اقرأ المزيد

مجموعة منتقاة من مصادر Generative UI وGenUI وواجهات AI الديناميكية.

صُنع باستخدام ❤️ من قِبل المجتمع