If your tool ran but the network dropped, PraisonAI tells you instead of running the tool twice.
Quick Start
1
Nothing to configure
Replay-safety is default-on for tool-calling turns. No new config keys, no new Agent params.
2
Handle the surfaced outcome
A surfaced
provider_outcome_unknown means the provider may already have run your tool — you choose whether to verify state, ask the user, or retry deliberately.The Invariant
On a tool-calling turn, a post-dispatch transient failure surfaces asprovider_outcome_unknown instead of auto-retrying. Tool-less turns and pre-dispatch failures still auto-retry exactly as before.
Replaying a post-dispatch failure on a tool-calling turn could double-execute the tool — double billing, or a re-run of a side-effecting action. The gate stops that.
Streaming → Non-Streaming Fallback is Gated Too
When streaming fails, PraisonAI used to silently retry the same request withstream=False. On a tool-calling turn that is still a provider replay — the tool may already have run server-side mid-response. The fallback now respects the same replay-safety gate.
Both fallback sites are gated:
Tool-less turns and pre-dispatch failures (connect / DNS / TLS) still fall back to non-streaming exactly as before.
Agent example
What Counts as Post-Dispatch (Replay-Unsafe)
Failures that happen once bytes are on the wire — the provider may already have produced output.
Message patterns treated as post-dispatch:
read timeout, connection reset, reset by peer, peer closed, incomplete read, chunked, response ended prematurely, server disconnected.
What Counts as Pre-Dispatch (Replay-Safe)
Failures that prove the request never reached the provider — always safe to replay, retried as before.
Message patterns treated as pre-dispatch:
connect timed out, connection refused, name resolution, getaddrinfo, handshake, certificate verify failed, cert verif.
The bare ssl and tls tokens are not pre-dispatch signals — ssl.SSLError is raised for both handshake failures and mid-response read/reset failures, so a transport label alone does not establish that the request failed before dispatch. Use the phase-specific handshake or certificate verify message to mark a TLS error as pre-dispatch.
Phase-Specific Precedence
A TLS handshake or certificate verification phase completes before any request byte is written, so an error naming either phase wins over a post-dispatch token.Ambiguous transient errors — plain 5xx or “service unavailable” — default to safe, preserving existing auto-retry behaviour.
Building Your Own Retry Wrapper
The classifier is public — import it if you write your own retry loop.is_replay_unsafe(error) returns True when the failure signals the provider may already have processed the request.
Walk the Exception Chain
The streaming path re-raises wrapping the cause (raise Exception(...) from read_timeout), so the replay-unsafe signal lives on __cause__ / __context__ rather than the outermost exception. Use is_replay_unsafe_chain() to walk the chain:
The chain is treated as unsafe if any link is unsafe — a pre-dispatch (safe) link does not override an unsafe one found elsewhere in the chain.
Common Patterns
A TLS-labelled read timeout or mid-response reset on a tool turn now surfaces instead of silently retrying.Best Practices
Treat provider_outcome_unknown as 'verify, then decide'
Treat provider_outcome_unknown as 'verify, then decide'
The provider may already have run your tool. Check downstream state — was the flight booked, the charge made — before retrying the turn.
Make side-effecting tools idempotent where you can
Make side-effecting tools idempotent where you can
Idempotency keys turn a risky replay into a safe one. When a tool is idempotent, a deliberate retry after
provider_outcome_unknown is cheap.Only block retries on tool-calling turns
Only block retries on tool-calling turns
Tool-less completions have no external side effects. Replaying a post-dispatch failure there is safe, so the gate only fires when the turn exposes tools.
Trust the classifier over message text
Trust the classifier over message text
is_replay_unsafe walks the exception’s class hierarchy first, so provider SDK subclasses are matched even when their messages vary. For paths that re-raise with from (streaming is one), use is_replay_unsafe_chain() so a wrapped cause is still detected.Related
Tools
Author side-effecting tools that behave well under replay-safety.
LLM Error Classification
How PraisonAI categorises provider errors for retry decisions.
How tool turns stay safe when a stream drops mid-response.

