StartupWatchdog guards the window between process entry and a live event loop, so a hanging import, a blocking credential mint, or a deadlocked connect() never becomes a silent zombie.
Running
praisonai gateway start? CLI/YAML wiring for this rung is not yet in the wrapper β it is deliberately scoped out of PR #5266 as an embedder concern. You need a small custom entry point (see Quick Start) to arm it today. Watch issue #5265 for the follow-up CLI/YAML PR.This primitive is a minimal, dependency-free core building block only β no Agent params, no new dependencies, no
praisonai gateway start --startup-watchdog-timeout=β¦ flag, and no gateway.startup_watchdog: YAML block. Wiring those into the wrapper is intentionally out of scope for PR #5266 and is left for a future embedder-side change. Arm it manually from your own entry point for now.Watchdog ladder
StartupWatchdog is the front rung of a three-rung ladder β it covers startup, LoopWatchdog covers the running loop and the shutdown drain.
Quick Start
1
Arm before heavy startup, confirm when the loop is live
Arm the deadline before the first heavy import, then disarm the instant the loop begins serving.
2
Extend the deadline across slow-boot phases
A legitimately slow boot reports progress at each phase, so it earns a fresh window instead of tripping.
3
Observe without exiting (dev / staging)
Use
on_expire="dump_only" to capture a stack dump on a wedge without killing the process.How It Works
A dedicated daemon OS thread counts down a wall-clock deadline, so it keeps working precisely when the main thread is wedged.
The watchdog is stdlib-only (
threading, time, os, sys, math, faulthandler) β it has to run when the code it guards cannot, so its correctness never depends on the imports it protects.
On expiry the stack-dump header reads:
75 contract.
Configuration Options
The constructor holds every knob;arm() takes the deadline.
arm(*, deadline_s: float, dump_stacks: bool = True):
StartupWatchdog exposes a small surface:
Common Patterns
Three ways teams wire the startup watchdog. Minimal embedder (production): arm before heavy startup, confirm when the loop is live, and wire the process to your existing supervisor.EX_TEMPFAIL (75) is the same code the rest of the restart-intent protocol uses.
dump_and_exit in production.
Fail-open guarantees
A misbehaving watchdog must never kill a healthy startup.arm()while already armed is a no-op.arm()withdeadline_s <= 0,NaN, orinfis a silent no-op β nothing to guard.report_startup_progress()when not armed is a no-op.confirm_loop_live()when not armed is safe.- A
confirm_loop_live()racing an in-flight expiry β including during the pre-os._exitstderr flush β suppresses the destructive exit, so a legitimate late confirm never terminates a healthy startup. - Any exception inside the watchdog is swallowed.
Best Practices
Arm before the first heavy import, not after
Arm before the first heavy import, not after
The watchdog guards imports too β a hanging SDK import is exactly the kind of pre-loop wedge it exists to catch. Arm it as the very first thing your entry point does.
Match the deadline to your 99th-percentile cold start
Match the deadline to your 99th-percentile cold start
Size
deadline_s off the slowest healthy boot you have seen, not the mean. A deadline tuned to the average trips on every cold cache warm-up.Report progress at every meaningful phase transition
Report progress at every meaningful phase transition
Config loaded, secrets minted, each adapter connected β one
report_startup_progress() per phase keeps a legitimately slow boot alive while a truly wedged phase still trips.Keep max_extensions honest
Keep max_extensions honest
A boot that always exhausts the extension budget is a boot that needs tuning, not a bigger budget. Raise the deadline for genuinely slow startups; keep
max_extensions bounded so a wedge inside one phase still fires.Confirm the loop live immediately before LoopWatchdog.arm(loop)
Confirm the loop live immediately before LoopWatchdog.arm(loop)
Disarm the startup rung the instant the running-loop rung takes over β never after. This closes the coverage handoff with no gap and no double-guard.
Wire your supervisor for EX_TEMPFAIL (75)
Wire your supervisor for EX_TEMPFAIL (75)
systemd (
Restart=on-failure), launchdβs KeepAlive, Dockerβs restart: unless-stopped, and Kubernetes all restart on exit 75 β the same code LoopWatchdog uses. See Exit Codes for the full contract.Related
Loop Watchdog
Rungs 2 & 3: probe a live loop, and guard the shutdown drain.
Exit Codes
The shared
EX_TEMPFAIL (75) restart-intent protocol.Forensics
What to keep on unhealthy exits.
Crash-Loop Guard
Guard against restart storms.

