ChipokiaTech news, without the noise
← All stories
InfoQ·

Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring

Google has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The update introduces long-horizon tasks (LHTs), agent-based evaluation, and continuous scoring to better assess…

Preview courtesy of InfoQ. The full article opens on their site.

More from InfoQ