Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
Researchers have created a new benchmark called MobileForge to evaluate the ability of AI models to generate complete mobile apps from visual designs. The benchmark assesses five key aspects: building, navigation, visual fidelity, code maintainability, and efficiency. Current multimodal large language models can create functional but incomplete apps that lack reliable interactive navigation. The study evaluates six frontier models using MobileForge and finds room for improvem
Researchers have created a new benchmark called MobileForge to evaluate the ability of AI models to generate complete mobile apps from visual designs. The benchmark assesses five key aspects: building, navigation, visual fidelity, code maintainability, and efficiency. Current multimodal large language models can create functional but incomplete apps that lack reliable interactive navigation. The study evaluates six frontier models using MobileForge and finds room for improvement in these areas.
---
Why it matters: This matters to AI engineers because it highlights the limitations of current design-to-code benchmarks and provides a new standard for evaluating the capabilities of multimodal large language models in generating complete mobile apps.
Source: https://arxiv.org/abs/2607.28645
This article was originally published at: https://arxiv.org/abs/2607.28645