Google built 'computer use' directly into Gemini 3.5 Flash. The agent sees the screen, reasons and clicks - in browser, on mobile, on desktop. Not a separate product. Part of the model.
- Google built computer use directly into Gemini 3.5 Flash (June 24): the agent sees, reasons and clicks on its own.
- Works in browser, mobile and desktop environments; aimed at long, multi-step tasks.
- Safeguards: training against misuse, confirmation on sensitive actions and prompt injection detection.
The agent looks at your screen and clicks instead of you. Google built 'computer use' directly into Gemini 3.5 Flash - it sees the interface, reasons and acts in browser, on mobile and on desktop.
Here I see the same line as with OpenAI and Anthropic - the model stops waiting to be asked and starts taking steps on its own. Google adds something important on top - safeguards for sensitive actions and protection against command injection.
The ability for an agent to click instead of you is the easy part. The hard part is putting brakes in the right places. That the two arrive together here, not patched on later, is why I'd run it on a real browser. With the others I'd still wait.
An agent that acts on its own on screen effectively needs an allowlist, confirmation on sensitive actions, logs and sandbox, not blind trust.