Sobre
Apple-affiliated research on evaluating whether generated interfaces support realistic user journeys. FlowEval runs validated tasks on reference websites and generated analogs, compares their navigation traces with reference-based similarity metrics, and tests how well those scores align with expert UI judgments.
Resumo
1. Interaction-first evaluation: Measures complete task flows instead of relying only on screenshots or code correctness
2. Trusted references: Compares generated sites with traces collected from known, high-quality interfaces
3. Automatable method: Uses computer-use agents and navigation tasks to collect comparable interaction sequences
4. Interpretable diagnostics: Trace similarity can reveal where a generated flow diverges from the reference experience
5. Human alignment evidence: Reports strong correlation between reference-based metrics and expert evaluations of generated UIs