Skip to main content
The gateway now ships in the praisonai-bot package. praisonai serve gateway still works exactly as documented here; for a standalone install see praisonai-bot Migration.
A scale-to-zero policy quiesces the gateway and (optionally) suspends the compute host when no turn is in flight and no background work is pending, then wakes on the next inbound message. The user is idle; the gateway scales workers to zero and wakes them when the next message arrives.

Quick Start

1

Simplest setup — 5-minute idle timeout

The gateway now ticks every 30 seconds. After 5 minutes of silence with no turns in flight, it stops transports and calls your _on_quiesce hook (if set). The next POST /_wake brings it back.
2

Tighter timeout for more aggressive savings

3

With a compute-host suspend hook (Fly / Modal)

_on_quiesce is set as an attribute after construction (not a constructor kwarg). Set it before calling start() or run(). When a public setter is added in a future release, this section will be updated.
4

Primary runtime — enable via gateway.yaml

Scale-to-zero is wired into the primary WebSocketGateway (praisonai gateway start) through a lifecycle: block.
The gateway serves normally, then quiesces after idle_minutes of no inbound or in-flight work. The next client frame wakes it.
5

Primary runtime — enable via CLI flags (no YAML)

The flags synthesize a lifecycle: block, so CLI overrides win over gateway.yaml. No wake_url is needed — the listening socket is an inherent self-wake path.

How It Works

Key behaviours:

Configuration Options

ScaleToZeroPolicy constructor

Properties on the instance:

BotOS — idle policy field

BotOS runtime activity hooks

Imports

These names export from praisonaiagents.gateway. Top-level praisonaiagents does not re-export them. BotOS itself imports from praisonai.bots.

lifecycle.scale_to_zero: YAML block (WebSocketGateway)

The primary WebSocketGateway reads scale-to-zero from a lifecycle: block at the top level of gateway.yaml or nested under gateway:.

CLI flags (WebSocketGateway)

CLI flags win over YAML — they are folded into the lifecycle: block before the policies are built.

Observability

When scale-to-zero (or a drain marker watch) is configured, health() adds a lifecycle object so operators can see the state without scraping logs.
The lifecycle field is emitted only when a lifecycle feature is configured, so always-on gateways are unchanged.

Hot Reload

Enabling or disabling scale-to-zero via a gateway.yaml reload takes effect without a full process restart. The gateway rebuilds the pure policy and cancels or relaunches only the idle loop whose enablement changed — an unchanged lifecycle: block leaves running tasks untouched. See Gateway Hot Reload.

Common Patterns

Fly Machines auto-suspend

The /_wake route in your web framework calls await botos.wake(), which reconnects all transports and the gateway is live again within seconds.

Env-var kill switch

Set SCALE_TO_ZERO=1 in production. In dev, leave it unset — the policy stays disabled and the gateway runs always-on with zero overhead.

Tighter timeout for prototype bots

Use 1 minute for disposable demo bots. Bump to 5+ minutes for production where users expect instant responses.

Best Practices

On BotOS, without a wake_url should_arm returns False and the gateway logs "idle policy not armed (no wake path); staying always-on" — the policy refuses to quiesce into a state it cannot recover from.On the primary WebSocketGateway (praisonai gateway start), the listening socket is treated as an inherent self-wake path: the gateway stays dormant with the socket open and wakes on the next inbound client frame. So wake_url is only required when you also set an on_quiesce host-suspend hook (Fly / Modal / Daytona), which powers the host down and needs an external endpoint to bring it back. CLI-only --scale-to-zero arms and quiesces correctly without a wake_url.
Humans pause between messages in the same conversation — typing, thinking, copy-pasting. A timeout under 2 minutes can spin transports down mid-conversation, adding a cold-start delay the user will feel. Start at 5 minutes for interactive bots and tune down once you have usage data.
Test that curl -X POST https://your-bot.example.com/_wake successfully brings the bot back before you wire _on_quiesce to actually suspend the host. That way you can recover from a broken suspend hook without needing to redeploy.
The policy already blocks dormancy whenever any enabled scheduled job exists. If you want true scale-to-zero, audit praisonai schedule list and disable jobs that should not keep the bot resident. Never rely on a long timeout to “outlast” a scheduled job — the job will prevent quiescing for as long as it is enabled.

Gateway Overview

The gateway architecture this feature layers on top of.

Session Continuity

Sessions survive the suspend — what makes resume after wake work.

BotOS

The multi-platform orchestrator that hosts the idle policy.

Drain Trigger

Sibling lifecycle policy — epoch-safe external drain marker.

Crash-Loop Guard

Sibling lifecycle policy — stop auto-resuming a crash-looping channel.