Acerca de
2026 benchmark for evaluating runnable LLM-generated web applications through coverage-guided browser exploration. IWC-Bench instruments each app, turns interaction traces into state-transition graphs, and scores visual aesthetics, usability, and requirement alignment across 369 real-world requirements and 5,088 acceptance criteria.
Resumen
1. Runtime evidence: Evaluates behavior reached through real browser interactions instead of crediting source code that may be unusable
2. Coverage-guided exploration: Uses instrumented code coverage to help an agent discover more implemented states and features
3. State-transition abstraction: Converts interaction traces into a structured graph before scoring the generated application
4. Separate quality dimensions: Measures visual aesthetics, usability, and requirement alignment rather than collapsing them into one score
5. Human-alignment evidence: Reaches 85.3% pairwise agreement with human preferences across 197 validated arena sessions