Skip to main content
Agents can read text from images - receipts, documents, signs, and screenshots.

Quick Start

1

Simple Usage

2

With Configuration


User Interaction Flow


Configuration Levels


Common Uses


API Reference

OCRConfig

Complete configuration options

OCRAgent

Full class documentation

Best Practices

Set detail: 'high' when reading receipts or documents.
“Extract the total and date” works better than “read this”.
Clear, well-lit images produce better results.

Vision

Image analysis

Files

File operations