Android Developers Blog
Follow
Evolving how LLMs are measured for Android: the next era of Android Bench
Android Bench, an LLM leaderboard for Android development tasks, was introduced in March to promote transparency and model improvements. Based on feedback, the benchmark now evaluates open-weight models and includes cost and efficiency dimensions. The July release adopts the Harbor framework, an updated benchmarking agent, for more rigorous model evaluations. This update necessitated re-running all models, causing a minor shift in scores, though historical data remains accessible. Eight new models, including Claude Fable 5, Claude Sonnet 5, and GLM 5.2, have been added to the leaderboard. Claude Fable 5 currently leads with an 84.5 score, while GLM 5.2 tops the open-weight models. The Android developer community can now contribute to Android Bench by submitting new development tasks or running and sharing benchmark evaluations. This initiative aims to create a benchmark that truly reflects the diverse challenges faced by global Android developers. The goal is to continuously update Android Bench to ensure AI assistance remains cutting-edge, smarter, and more effective. Developers are encouraged to engage with the GitHub repository and Harbor Hub for contributions and to review the updated leaderboard and methodology.