概要
August 2026 benchmark for evaluating LLM-generated interfaces with deterministic annotations instead of noisy human labels. SchemaGUI synthesizes bilingual instructions and function-call references from parameterized UI schemas, then measures feasibility, geometry, and layout complexity across six scenarios and five models.
要約
1. Deterministic evaluation: Generates exact function-call references from interface schemas instead of relying on subjective screenshot labels
2. Bilingual coverage: Evaluates 1,000 instances per scenario and language across six representative English and Chinese GUI tasks
3. Separates capability dimensions: Scores schema feasibility, geometric control, and overall GUI quality rather than collapsing them into one visual metric
4. Exposes spatial bottlenecks: Shows that model scaling improves valid structure faster than precise coordinates, especially in dense grids and multi-region layouts
5. Useful inference finding: Reports that thinking mode consumes more tokens while generally lowering GUI scores, particularly for smaller models