Google Integrates Computer Control Capabilities into Gemini 3.5 Flash
Google has directly integrated the "Computer Use" capability into its AI model "Gemini 3.5 Flash," enabling autonomous control of computers, browsers, and mobile devices. The model achieved a score of 78.4 on the "OSWorld" benchmark for evaluating PC operation tasks, which is considered equivalent to GPT-4.5 performance. Developers can access this functionality through the Gemini API, with anticipated applications in software testing automation and office task automation.

Google has announced the direct integration of "Computer Use" functionality into its AI model "Gemini 3.5 Flash." This enables the model to autonomously control personal computers, browsers, and mobile devices while recognizing their screens.
"Computer Use" refers to the capability for AI to operate computers on behalf of humans. For example, AI can independently judge and execute tasks such as opening a browser to fill in forms or launching applications to perform specific operations in sequence. Previously, such capabilities required dedicated software or complex configuration, but now this capability is embedded directly within the model itself.
The approach of AI autonomously controlling computers is a field that has recently garnered significant industry-wide attention as "AI agents." Major AI companies including OpenAI and Anthropic are each developing similar capabilities, and Google's direct integration of this function into its model represents a strategic move within this competitive landscape.
In terms of performance, the model achieved a score of 78.4 on the "OSWorld" benchmark, which is widely used to evaluate PC operation tasks. This score is considered equivalent to OpenAI's GPT-4.5 performance. OSWorld is a benchmark that measures how accurately AI can perform operations on actual operating systems, with higher scores indicating greater reliability in practical scenarios.
Developers can access this functionality through Google's API (application interface connecting different programs) called "Gemini API." Anticipated use cases include software testing automation and office task automation, with expectations for application in developing tools that enhance corporate operational efficiency.
The significance of this functionality integration extends beyond mere performance improvement. With AI now capable of completing tasks from screen recognition to operation execution as a standalone model, there is an emerging environment where AI agents can assume many routine PC tasks that previously required human intervention. If developers can more easily construct agents, the pace of enterprise system integration could accelerate significantly.
The key point to monitor going forward is how stably the system performs in actual operational environments. Since benchmark metrics do not necessarily correlate with real-world reliability, validation under diverse conditions becomes critical. The provision format through Gemini API sets up an environment where many developers can easily experiment, which could serve as a foundation for broader adoption.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.