Skip to main content
--max-tokens caps the model’s output length. When you pass it explicitly, PraisonAI honours it verbatim — even when the value equals the built-in default.

Quick Start

1

Set a budget from your agent

2

Cap output from the CLI

Behaviour change (PR #5360 / issue #5358). Before this fix, --max-tokens 16000 was silently dropped because the resolver could not distinguish “user typed 16000” from “argparse filled in the default”. The parser now marks explicitness with a sentinel at parse time, so --max-tokens 16000 behaves identically to --max-tokens 15999. Any prior workaround like --max-tokens 16001 is no longer needed.

How It Works

The parser tracks whether you actually typed the flag, so an explicit value always wins — regardless of what it equals.

Precedence

The most specific source that you actually set wins.

Usage


Common Patterns

Short Response

Long-form Content

With Research


Token Limits by Model

Setting max-tokens higher than the model’s limit will be capped automatically.

Best Practices

An explicit --max-tokens N always wins — including --max-tokens 16000. You no longer need workarounds like --max-tokens 16001.
Omit --max-tokens to let each agent’s max_tokens: in YAML apply. The CLI default only fills in when nothing else set a budget.
Values above the model’s output limit are capped automatically, so match your budget to the model you run.

Agent Max Budget

Set per-agent token and cost budgets in code.

Model-aware clamp (#4124)

Tracking issue for clamping budgets to each model’s cap.