О материале
September 2026 benchmark connecting aesthetic scoring, diagnosis, repair, and text-to-UI generation across 1,395 executable web interfaces. It pairs diagnosis and repair on 660 controlled-degradation cases to test whether models can translate design judgments into effective interface changes.
Краткое содержание
1. Shared evaluation foundation: Tests four capabilities against the same interface pool and professional aesthetic principles
2. Judgment-to-action evidence: Aligns defect diagnosis with source-code repair on the same degraded page
3. Measured capability gap: Across 12 models, exact diagnosis-chain success peaks at 24.7%, and correct judgments do not consistently lead to successful repairs
4. Reproducibility boundary: The dataset is public, while the linked repository still lists evaluation code as forthcoming