Skip to main content
Tasks with quality_check=True score outputs and store high-quality results in memory.
The user runs a team task; quality scoring decides whether outputs land in long-term memory.

How It Works

Quick Start

1

Simple Usage

Enable quality checking on a task (default is True):
2

With Configuration

Disable for fast runs or use execution presets:

How It Works

When quality_check=True and memory is configured:
  1. Agent completes the task
  2. Memory.calculate_quality_metrics() scores completeness, relevance, clarity, accuracy via LLM
  3. finalize_task_output() stores in long-term memory only when score exceeds 0.7
  4. Quality metadata attaches to the task result
Memory is required — without it, quality checking logs a warning and skips storage.

Endpoint and Key Resolution

Memory.calculate_quality_metrics() picks its scoring endpoint and key from the memory config, so quality scoring reaches the same host you point memory at. Config keys are read from the top level or a nested "config" dict, e.g. Memory(config={"base_url": ..., "api_key": ...}).
The litellm path now honours config base_url / api_key (passed as litellm’s api_base kwarg). Before PR #4958 it silently went to OpenAI’s default endpoint — users pointing memory at Groq, OpenRouter, Together, or a private OpenAI-compatible endpoint will now see the quality-scoring call go to the configured host.

Configuration Options

Execution presets: "fast" disables quality check; "balanced" and "thorough" enable it.

Best Practices

Clear expectations produce meaningful scores — vague tasks score inconsistently.
Quality checking stores to memory — attach memory=Memory() to the agent or task.
Set quality_check=False on brainstorming or speed-critical tasks.
Search with min_quality=0.7 to reuse past strong outputs as context.

Quality-Based RAG

Quality scoring for retrieval

Memory

Memory configuration and search