Understand images, videos, documents, and charts with Qwen.
visual analysis
Analyzes images and videos for understanding, reasoning, comparison, and text extraction.
When to use it
Use for image or video analysis, OCR, chart and table reading, visual reasoning, comparisons, and screenshot understanding.
Give it visual input and the desired analysis; it saves a response containing the resulting understanding, reasoning, or extracted text.
What you provide
This skill
DashScope API
Sends visual input to DashScope
DashScope temp storage
Uploads large files temporarily
OSS bucket
Uploads files to your OSS
Qwen vision model catalog
Reads the current model catalog
The default script path requires Python 3.9 or newer and uses only the standard library.
Reads the required API key from the environment and can also load it from .env.
Reads QIANWEN_API_KEY as an alternative environment-variable alias for DASHSCOPE_API_KEY.
This package enables the documented production path that uploads files to the user's own OSS bucket.