Назад ΠΊ списку
πŸ“„Papers

FlowEval: Reference-Based Evaluation of Generated User Interfaces

Wu J. et al.
2026-05

О ΠΌΠ°Ρ‚Π΅Ρ€ΠΈΠ°Π»Π΅

Apple-affiliated research on evaluating whether generated interfaces support realistic user journeys. FlowEval runs validated tasks on reference websites and generated analogs, compares their navigation traces with reference-based similarity metrics, and tests how well those scores align with expert UI judgments.

ΠšΡ€Π°Ρ‚ΠΊΠΎΠ΅ содСрТаниС

Why FlowEval Adds New Signal:

1. Interaction-first evaluation: Measures complete task flows instead of relying only on screenshots or code correctness
2. Trusted references: Compares generated sites with traces collected from known, high-quality interfaces
3. Automatable method: Uses computer-use agents and navigation tasks to collect comparable interaction sequences
4. Interpretable diagnostics: Trace similarity can reveal where a generated flow diverges from the reference experience
5. Human alignment evidence: Reports strong correlation between reference-based metrics and expert evaluations of generated UIs

Π’Π΅Π³ΠΈ

ui-evaluationinteraction-flowscomputer-use-agentsbenchmarkhuman-evaluation
Π§ΠΈΡ‚Π°Ρ‚ΡŒ ΠΎΡ€ΠΈΠ³ΠΈΠ½Π°Π»

ΠŸΠΎΡ…ΠΎΠΆΠΈΠ΅ ΠΌΠ°Ρ‚Π΅Ρ€ΠΈΠ°Π»Ρ‹

πŸ“„Papers

LEGOUI: Designing with UI-DSL Bricks for Transparent, Controllable Generation

August 2026 HCI paper introducing a staged Generative UI framework for early design ideation. LEGOUI records prompt-derived and model-inferred decisions in a provenance-aware UI DSL, lets users accept, reject, or add decisions across layout, interaction, relation, and style stages, and renders from the evolving specification.

genuiui-dsldesign-time
ΠΎΡ‚ Zhou Y. et al. Β· 2026-08
Π§ΠΈΡ‚Π°Ρ‚ΡŒ дальшС
πŸ“„Papers

Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation

July 2026 paper introducing ESPP, an evaluation method that replaces one LLM judge with a panel of evidence-grounded, psychologically diverse personas. Panelists independently rate generated-interface screenshots, exchange opinions through a bounded-confidence mechanism, and are combined with social weighting.

genuiui-evaluationhuman-preferences
ΠΎΡ‚ Wu Z. et al. Β· 2026-07
Π§ΠΈΡ‚Π°Ρ‚ΡŒ дальшС
πŸ“„Papers

Design Theater: A Benchmark for Generative UI

July 2026 paper introducing a benchmark for checking whether Generative UI tools implement the design rationales they present to users. Across 120 interfaces from five tools and 24 structural, styling, and functional tasks, the study measures gaps between claimed design decisions and the generated result.

genuiui-evaluationdesign-rationale
ΠΎΡ‚ Imteyaz K. et al. Β· 2026-07
Π§ΠΈΡ‚Π°Ρ‚ΡŒ дальшС

ΠŸΠΎΠ΄Π±ΠΎΡ€ΠΊΠ° рСсурсов ΠΏΠΎ Generative UI, GenUI ΠΈ Π³Π΅Π½Π΅Ρ€Π°Ρ‚ΠΈΠ²Π½Ρ‹ΠΌ интСрфСйсам.

Π‘Π΄Π΅Π»Π°Π½ΠΎ с ❀️ сообщСством