Skip to main content
StartupWatchdog guards the window between process entry and a live event loop, so a hanging import, a blocking credential mint, or a deadlocked connect() never becomes a silent zombie.
Running praisonai gateway start? CLI/YAML wiring for this rung is not yet in the wrapper β€” it is deliberately scoped out of PR #5266 as an embedder concern. You need a small custom entry point (see Quick Start) to arm it today. Watch issue #5265 for the follow-up CLI/YAML PR.
This primitive is a minimal, dependency-free core building block only β€” no Agent params, no new dependencies, no praisonai gateway start --startup-watchdog-timeout=… flag, and no gateway.startup_watchdog: YAML block. Wiring those into the wrapper is intentionally out of scope for PR #5266 and is left for a future embedder-side change. Arm it manually from your own entry point for now.

Watchdog ladder

StartupWatchdog is the front rung of a three-rung ladder β€” it covers startup, LoopWatchdog covers the running loop and the shutdown drain.

Quick Start

1

Arm before heavy startup, confirm when the loop is live

Arm the deadline before the first heavy import, then disarm the instant the loop begins serving.
2

Extend the deadline across slow-boot phases

A legitimately slow boot reports progress at each phase, so it earns a fresh window instead of tripping.
3

Observe without exiting (dev / staging)

Use on_expire="dump_only" to capture a stack dump on a wedge without killing the process.

How It Works

A dedicated daemon OS thread counts down a wall-clock deadline, so it keeps working precisely when the main thread is wedged. The watchdog is stdlib-only (threading, time, os, sys, math, faulthandler) β€” it has to run when the code it guards cannot, so its correctness never depends on the imports it protects. On expiry the stack-dump header reads:
This is distinct from both the loop-wedge header and the shutdown-phase header, so a single glance at the log tells you which rung fired. See Exit Codes for the shared 75 contract.

Configuration Options

The constructor holds every knob; arm() takes the deadline. arm(*, deadline_s: float, dump_stacks: bool = True): StartupWatchdog exposes a small surface:

Common Patterns

Three ways teams wire the startup watchdog. Minimal embedder (production): arm before heavy startup, confirm when the loop is live, and wire the process to your existing supervisor. EX_TEMPFAIL (75) is the same code the rest of the restart-intent protocol uses.
Slow boot with progress reports: many adapters or cold DNS make a healthy boot take longer than a single window β€” report progress at every phase so the deadline extends.
Dump-only staging: catch a hang without killing the process, inspect the dump, then flip to dump_and_exit in production.

Fail-open guarantees

A misbehaving watchdog must never kill a healthy startup.
  • arm() while already armed is a no-op.
  • arm() with deadline_s <= 0, NaN, or inf is a silent no-op β€” nothing to guard.
  • report_startup_progress() when not armed is a no-op.
  • confirm_loop_live() when not armed is safe.
  • A confirm_loop_live() racing an in-flight expiry β€” including during the pre-os._exit stderr flush β€” suppresses the destructive exit, so a legitimate late confirm never terminates a healthy startup.
  • Any exception inside the watchdog is swallowed.

Best Practices

The watchdog guards imports too β€” a hanging SDK import is exactly the kind of pre-loop wedge it exists to catch. Arm it as the very first thing your entry point does.
Size deadline_s off the slowest healthy boot you have seen, not the mean. A deadline tuned to the average trips on every cold cache warm-up.
Config loaded, secrets minted, each adapter connected β€” one report_startup_progress() per phase keeps a legitimately slow boot alive while a truly wedged phase still trips.
A boot that always exhausts the extension budget is a boot that needs tuning, not a bigger budget. Raise the deadline for genuinely slow startups; keep max_extensions bounded so a wedge inside one phase still fires.
Disarm the startup rung the instant the running-loop rung takes over β€” never after. This closes the coverage handoff with no gap and no double-guard.
systemd (Restart=on-failure), launchd’s KeepAlive, Docker’s restart: unless-stopped, and Kubernetes all restart on exit 75 β€” the same code LoopWatchdog uses. See Exit Codes for the full contract.

Loop Watchdog

Rungs 2 & 3: probe a live loop, and guard the shutdown drain.

Exit Codes

The shared EX_TEMPFAIL (75) restart-intent protocol.

Forensics

What to keep on unhealthy exits.

Crash-Loop Guard

Guard against restart storms.