Turn images into answers you can use.
image understanding
Delegates image understanding to a configurable OpenAI-compatible vision model and returns a text answer.
When to use it
Use it whenever a task requires understanding image content that the agent cannot view directly, including OCR, screenshots, diagrams, charts, and scanned documents.
Give it an image path or URL and a clear question; it returns the vision model's text answer.
What you provide
This skill
OpenAI-compatible vision endpoint
Sends images to vision endpoint
.claude/settings.json files
Reads vision configuration files
OpenClaw dotenv files
Reads OpenClaw configuration
Python 3 must be available to run the script.
A configured OpenAI-compatible vision model endpoint must be available; the base URL and model are required.
An API key can be supplied through the environment or settings configuration; it is optional for local servers.