Skip to main content
Agents can see and understand images - describe content, read text, and answer questions.

Quick Start

1

Simple Usage

2

With Configuration


User Interaction Flow


Configuration Levels


What You Can Do


API Reference

VisionConfig

Complete configuration options

VisionAgent

Full class documentation

Best Practices

Use GPT-4o, Claude 3, or Gemini Pro Vision for image analysis.
“What text is on the document?” works better than “What is this?”
Set detail: 'high' when reading small text or documents.

Video

Analyze videos

OCR

Extract text from images