ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
4
Citations
0
Influential Citations
arXiv.org
Venue
2025
Year
With the advancements in hardware, software, and large language model technologies, the interaction between humans and operating systems has evolved from the command-line interface to the rapidly emerging AI agent interactions. Building an operating system (OS) agent capable of executing user instructions and faithfully following user desires is becoming a reality. In this technical report, we present ColorAgent, an OS agent designed to engage in long-horizon, robust interactions with the environment while also enabling personalized and proactive user interaction. To enable long-horizon interactions with the environment, we enhance the model's capabilities through step-wise reinforcement learning and self-evolving training, while also developing a tailored multi-agent framework that ensures generality, consistency, and robustness. In terms of user interaction, we explore personalized user intent recognition and proactive engagement, positioning the OS agent not merely as an automation tool but as a warm, collaborative partner. We evaluate ColorAgent on the AndroidWorld and AndroidLab benchmarks, achieving success rates of 77.2% and 50.7%, respectively, establishing a new state of the art. Nonetheless, we note that current benchmarks are insufficient for a comprehensive evaluation of OS agents and propose further exploring directions in future work, particularly in the areas of evaluation paradigms, agent collaboration, and security.
ColorAgent addresses a critical gap in the development of operating system (OS) agents: the ability to handle long-horizon, robust interactions while also providing personalized and proactive user engagement. As human-OS interaction shifts from command-line interfaces to AI agent interactions, the need for agents that can faithfully execute user instructions over extended periods becomes paramount. This paper introduces a comprehensive approach that combines reinforcement learning, self-evolving training, and a multi-agent framework to achieve state-of-the-art performance on two major Android benchmarks.
The significance extends beyond raw performance. By emphasizing personalized intent recognition and proactive engagement, ColorAgent repositions OS agents from simple automation tools to collaborative partners. This aligns with the broader trend in AI toward more human-centric and interactive systems, making the paper relevant not only to reinforcement learning researchers but also to those working on human-computer interaction, personalization, and multi-agent systems.
ColorAgent achieves a success rate of 77.2% on AndroidWorld and 50.7% on AndroidLab, both establishing new state-of-the-art results. These benchmarks are widely used for evaluating OS agents, and the improvements over prior methods are significant. The paper does not provide direct comparisons in the abstract, but the success rates indicate a substantial leap in capability, as previous state-of-the-art results were lower.
The broader impact of ColorAgent lies in its demonstration that OS agents can be both high-performing and user-centric. The combination of reinforcement learning and multi-agent collaboration sets a new standard for building robust, long-horizon agents. Moreover, the focus on personalization and proactive engagement opens new avenues for research in human-AI collaboration. The paper also highlights the inadequacy of current benchmarks, urging the community to develop more comprehensive evaluation paradigms that consider collaboration, security, and user experience. This work will likely influence future OS agent designs and benchmark development, pushing the field toward more holistic and practical AI systems.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba