Skip to main content
On a tool-calling turn, PraisonAI never silently retries a failure that happened after the request was dispatched — not from the retry driver, and not from the streaming → non-streaming fallback. The post-dispatch error surfaces to your code so you decide what to do.
If your tool ran but the network dropped, PraisonAI tells you instead of running the tool twice.

Quick Start

1

Nothing to configure

Replay-safety is default-on for tool-calling turns. No new config keys, no new Agent params.
2

Handle the surfaced outcome

A surfaced provider_outcome_unknown means the provider may already have run your tool — you choose whether to verify state, ask the user, or retry deliberately.

The Invariant

On a tool-calling turn, a post-dispatch transient failure surfaces as provider_outcome_unknown instead of auto-retrying. Tool-less turns and pre-dispatch failures still auto-retry exactly as before. Replaying a post-dispatch failure on a tool-calling turn could double-execute the tool — double billing, or a re-run of a side-effecting action. The gate stops that.
Extended in PR #5443 — generic TLS/SSL labels no longer mask explicit post-dispatch signals. The replay-safety contract was first introduced for #3860.

Streaming → Non-Streaming Fallback is Gated Too

When streaming fails, PraisonAI used to silently retry the same request with stream=False. On a tool-calling turn that is still a provider replay — the tool may already have run server-side mid-response. The fallback now respects the same replay-safety gate. Both fallback sites are gated: Tool-less turns and pre-dispatch failures (connect / DNS / TLS) still fall back to non-streaming exactly as before.

Agent example


What Counts as Post-Dispatch (Replay-Unsafe)

Failures that happen once bytes are on the wire — the provider may already have produced output. Message patterns treated as post-dispatch: read timeout, connection reset, reset by peer, peer closed, incomplete read, chunked, response ended prematurely, server disconnected.

What Counts as Pre-Dispatch (Replay-Safe)

Failures that prove the request never reached the provider — always safe to replay, retried as before. Message patterns treated as pre-dispatch: connect timed out, connection refused, name resolution, getaddrinfo, handshake, certificate verify failed, cert verif. The bare ssl and tls tokens are not pre-dispatch signals — ssl.SSLError is raised for both handshake failures and mid-response read/reset failures, so a transport label alone does not establish that the request failed before dispatch. Use the phase-specific handshake or certificate verify message to mark a TLS error as pre-dispatch.

Phase-Specific Precedence

A TLS handshake or certificate verification phase completes before any request byte is written, so an error naming either phase wins over a post-dispatch token.
Ambiguous transient errors — plain 5xx or “service unavailable” — default to safe, preserving existing auto-retry behaviour.

Building Your Own Retry Wrapper

The classifier is public — import it if you write your own retry loop.
is_replay_unsafe(error) returns True when the failure signals the provider may already have processed the request.

Walk the Exception Chain

The streaming path re-raises wrapping the cause (raise Exception(...) from read_timeout), so the replay-unsafe signal lives on __cause__ / __context__ rather than the outermost exception. Use is_replay_unsafe_chain() to walk the chain:
The chain is treated as unsafe if any link is unsafe — a pre-dispatch (safe) link does not override an unsafe one found elsewhere in the chain.

Common Patterns

A TLS-labelled read timeout or mid-response reset on a tool turn now surfaces instead of silently retrying.

Best Practices

The provider may already have run your tool. Check downstream state — was the flight booked, the charge made — before retrying the turn.
Idempotency keys turn a risky replay into a safe one. When a tool is idempotent, a deliberate retry after provider_outcome_unknown is cheap.
Tool-less completions have no external side effects. Replaying a post-dispatch failure there is safe, so the gate only fires when the turn exposes tools.
is_replay_unsafe walks the exception’s class hierarchy first, so provider SDK subclasses are matched even when their messages vary. For paths that re-raise with from (streaming is one), use is_replay_unsafe_chain() so a wrapped cause is still detected.

Tools

Author side-effecting tools that behave well under replay-safety.

LLM Error Classification

How PraisonAI categorises provider errors for retry decisions.
How tool turns stay safe when a stream drops mid-response.