Automate browser tasks with model-guided control.
browser automation
Automates browser tasks with Gemini Computer Use and Playwright, returning the agent's result.
When to use it
Use it to automate browser tasks, build the Gemini Computer Use agent loop, or add safety confirmation for risky UI actions.
Give it a browser task prompt and starting URL; it runs the agent loop and prints the agent output.
What you provide
This skill
Gemini API
Sends prompts and screenshots
web pages
Reads requested web pages
Chromium
Launches Chromium browser
Gemini Computer Use
Gemini agent loop
Python must be available to create the virtual environment and run the agent script.
The google-genai package must be installed for Gemini API access.
The Playwright package must be installed for browser control.
The Gemini API key must be exported in GEMINI_API_KEY before running.