Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring
Google has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The update introduces long-horizon tasks (LHTs), agent-based evaluation, and continuous scoring to better assess…
Adjust this page to suit you. Your choices are saved in this browser.
Cookies. Chipokia uses one cookie, and only if you allow it: a random ID so your up and down votes stay yours. No analytics, no advertising, no tracking of any kind. Cookie & Privacy Policy.