Google's AI Models Face Tough New Test in Android Development
Google has unveiled an updated benchmark for evaluating how well large language models (LLMs) and AI agents handle complex Android development tasks. The new version, called Android Bench 2.0, includes more challenging tasks that can take days to complete, such as upgrading dependencies, adding major features, and building Android apps from scratch.
The benchmark uses a 'continuous scoring' system, which provides a more nuanced indication of how well a model performed, even if it didn't fully complete a task. Google has already tested several AI models using the new benchmark, with GPT-6 Astra leading the pack at 28% pass rate.
Gemini 3.8 Flash scored significantly lower, at just 8%. The results are intended to help developers understand which models are best suited for different Android development tasks. Google plans to continue expanding the leaderboard with more models and results over time.