рдкрд░рд┐рдЪрдп
August 2026 benchmark testing whether LLM assistants can maintain one executable web interface as user requirements evolve. EvoGenUI-Bench contains 150 five-turn tasks across information presentation, stateful interaction, and tool-grounded external state, with browser-based evaluation over screenshots, DOM and source evidence, actor traces, and runtime logs.
рд╕рд╛рд░рд╛рдВрд╢
1. Evaluates interface evolution: Tests cumulative revisions to one artifact instead of isolated page generation
2. Measures requirement retention: Introduces Adjacent Pass Retention to show whether a correct interface stays correct after the next request
3. Executes every artifact: Grounds scoring in browser behavior, rendered output, source, DOM, traces, logs, and external-state readback
4. Exposes a reliability gap: The strongest model reaches 74.9% turn success but completes only 37.3% of five-turn episodes
5. Diagnoses concrete failures: Separates information-architecture, derived-state, affordance-binding, external-grounding, and requirement-decomposition breakdowns