The gateway now ships in the
praisonai-bot package. praisonai serve gateway still works exactly as documented here; for a standalone install see praisonai-bot Migration.GatewayClient is a reconnecting WebSocket client that handles version negotiation, exponential backoff with jitter, and event-sequence gap detection — so your integration stays connected without you writing the socket loop.
GatewayClient; the client negotiates the handshake, streams events, and reconnects after disconnects.
Quick Start
1
Simplest Usage
2
With Backoff Tuning
3
With Gap and State Callbacks
How It Works
Calling
client.connect() on an instance whose previous loop had a server-supplied retry_after floor clears that floor before the new loop starts. Reusing a GatewayClient after a terminal failure will not inherit a stale backoff from the prior connection cycle — the clear happens on every connect() call, not only the first.Configuration Options
GatewayClient constructor:
BackoffConfig fields:
Connection states (from
ConnectionState):
Callbacks (set as attributes on the client instance):
send() method:
Returns:
Optional[str] — the request_id used (auto-generated when needed), or None when no idempotency key was involved.
Protocol Version Negotiation
The client and server negotiate a protocol version during thejoin handshake.
- Constants in
protocols.py:PROTOCOL_VERSION = 1,MIN_PROTOCOL_VERSION = 1,MAX_PROTOCOL_VERSION = 1. - Client sends
min_versionandmax_versionin thejoinmessage. - Server replies in
joinedwithprotocol_version,server_min_version,server_max_version.
min_version/max_version fields (non-integer, or min > max) produce code: "invalid_protocol_hello" and raise ConnectionError.
Terminal vs Transient Connect Errors
Not every connect failure should be retried. The client uses a single classifier —is_recoverable(code) from praisonaiagents.gateway.protocols — to decide whether to back off and try again, or stop the reconnect loop and surface the failure through on_reconnect_paused.
agent_not_found is terminal because the gateway emits it with a do_not_retry recovery step — the agent id in your hello frame does not exist on the server, so retrying will not help until a human fixes the id or registers the agent. The client used to loop forever on this code; it now pauses immediately.Gap Detection
Every event carries a monotonicsequence field; the client tracks _expected_sequence and fires on_gap when there’s a mismatch.
client.resync() resets the cursor to 0 and reconnects, triggering a full state reload from the server.
Projected State (De-duplicated View)
client.project() folds the raw event stream into a correct, resumable session view — no manual de-dup, gap-healing or bounded retention.
project(snapshot=None, *, max_tracked_runs=200) seeds from an optional snapshot, then yields a SessionProjectionState after every event. See Session Projection for the full behaviour, config options, and de-dup semantics.
Resume Snapshot
When reconnecting with a storedsession_id and since=cursor, the server replies with a single joined payload that restores full client state in one round trip.
The joined payload includes:
This means a reconnecting client learns current presence and health without extra requests — one round trip restores all state.
Common Patterns
- Basic Reconnecting Consumer
- Gap Handler with Resync
- Capped Retries
- Authenticated
- Branch on Streaming Events
Network Blip: What Users See
When a network blip occurs, the client handles reconnection silently — users see no errors and events resume from the cursor automatically. The bot keeps responding because:- Events are buffered until the cursor is confirmed
- The
since=cursoron reconnect replays any events delivered while offline - Users see continuous responses with no error messages
Exactly-Once Send (Queue and Flush)
One flag makes an outbound message survive a reconnect and run exactly once.request_id, an in-flight send() is either lost on a socket drop (at-most-once) or double-runs the turn if you resend (at-least-once). Attaching a stable request_id lets the server dedup a resend, so a queued-and-flushed message runs exactly once.
What the user experiences
A user chatting through the bot sends a message during a Wi-Fi blip; withqueue=True the message is held, flushed on reconnect, and the agent runs the turn once.
How it works
Choose your mode
Return value and correlation
send() returns Optional[str] — the request_id actually used (auto-generated when needed), or None when no idempotency key was involved. Persist it to correlate the accepted ack and match the eventual final response.
Semantic guarantees (pinned by the request-idempotency test suite):
- Same
request_id→ the server dedups (the turn runs once). - Distinct
request_ids → both turns run. - Missing
request_id→ legacy behaviour (no dedup — keeps existing clients working). - If the turn never actually ran (agent unavailable,
/stop,/cancel, or the enqueue itself raised), the reservation is released so a later retry with the same id gets a real chance.
Common patterns
- Queue while disconnected
- Caller-supplied stable id
- Legacy fire-and-forget
Best Practices
Pick a backoff that matches your network
Pick a backoff that matches your network
The defaults (
initial=1.0, max=30.0, multiplier=2.0) suit residential and cellular networks. For LAN-only deployments with fast recovery, tighten initial to 0.1 and max to 5.0.Handle terminal failures with on_reconnect_paused
Handle terminal failures with on_reconnect_paused
The reconnect loop stops on terminal codes (
auth_*, pairing_required, protocol_unsupported, agent_not_found, origin_not_allowed, configuration_error). Without a callback you have no way to know it stopped — logs are the only signal. Always set on_reconnect_paused before connect() and route the code to re-auth, re-pair, or an ops alert. See Terminal vs Transient Connect Errors for the full classification.Treat version_unsupported as a permanent failure
Treat version_unsupported as a permanent failure
connect() raises ValueError on version_unsupported and stops retrying. Wrap it and alert your ops team — this means the server and client are incompatible and need a coordinated upgrade. This is the protocol-negotiation counterpart to the other terminal codes covered in Terminal vs Transient Connect Errors.Wire on_gap to your replay or snapshot path
Wire on_gap to your replay or snapshot path
If your app already has its own snapshot mechanism, prefer it over
resync(). A targeted snapshot is faster than a full cursor reset.Pin max_reconnect_attempts in batch jobs
Pin max_reconnect_attempts in batch jobs
Long-running batch jobs should not retry forever on a dead gateway. Set
max_reconnect_attempts so the job fails fast and can be requeued.Related
Gateway
WebSocket control plane for multi-agent coordination
Gateway Overview
Architecture and deployment patterns
Session Persistence
How sessions survive across processes
Stream Events
Live progress events forwarded over WebSocket
Session Projection
De-duplicated, gap-aware view via
client.project()
