Quick Start
1
Simple Usage
2
With Configuration
User Interaction Flow
Configuration Levels
What You Can Do
API Reference
VisionConfig
Complete configuration options
VisionAgent
Full class documentation
Best Practices
Use vision-capable models
Use vision-capable models
Use GPT-4o, Claude 3, or Gemini Pro Vision for image analysis.
Be specific in questions
Be specific in questions
“What text is on the document?” works better than “What is this?”
Use high detail for text
Use high detail for text
Set
detail: 'high' when reading small text or documents.Related
Video
Analyze videos
OCR
Extract text from images

