Skip to main content
The user calls a hardened recipe HTTP endpoint; rate limits, metrics, and auth guard production traffic. Harden a recipe HTTP server with rate limits, Prometheus metrics, admin reload, and distributed tracing.

How It Works

Quick Start

1

Simple Usage

Enable rate limiting and metrics on an existing server:
Scrape metrics at GET /metrics (Prometheus format).
2

With Configuration

Production setup with auth, admin reload, workers, and tracing:

How It Works

Advanced features layer onto the base recipe server from Recipe Serve. Rate limiting uses a sliding window per client; metrics expose request counts and latency; admin endpoints hot-reload recipes without restart.

Rate limiter internals

The middleware buckets requests by the strongest identity the auth layer has already validated — never by a raw client header — so rotating a header cannot defeat the limit. RateLimiter exposes two entry points:
  • check(client_id) — synchronous, thread-safe (backward-compatible). check_sync remains as an alias.
  • check_async(client_id) — async, guarded by an asyncio.Lock. Used by the middleware hot path so overlapping requests can never both pass the length check before either appends.
If you rely on rate limiting in auth: none or auth: jwt mode, the bucket id is derived from the validated identity (JWT sub or client IP), never from the raw X-API-Key header. A client cannot rotate X-API-Key values to escape the limit. In auth: api-key mode the validated key is the bucket, so distinct valid keys naturally get distinct buckets — this is intentional.
Precedence: jwt:<sub> > apikey:<key> (only in api-key mode) > ip:<host> > "anonymous".

Configuration Options

rate_limit is enforced per worker process. When workers > 1, serve() splits the configured value across workers (rate_limit // workers, minimum 1) and emits a UserWarning at startup. This is best-effort — for exact cross-process limiting, use a shared store.

Common Patterns

Programmatic rate limiter

Admin reload

Prometheus scrape config

OpenTelemetry dependencies

OpenTelemetry is lazily imported — if packages are missing, the server logs a warning and continues without tracing.

Best Practices

Set enable_admin=True only alongside auth: api-key and load the key from PRAISONAI_API_KEY. Admin reload can change live behaviour — protect it.
Set workers to roughly 2 × CPU cores + 1. Workers above 1 disable hot reload automatically. Under the hood, serve(workers=N>1) hands uvicorn an app-factory import string (praisonai.recipe.serve:_app_factory) and ships the config to each worker through the PRAISONAI_RECIPE_SERVE_CONFIG environment variable. Previously, passing workers=N silently degraded to a single worker — that regression is fixed.
Keep /health and /metrics in rate_limit_exempt_paths so orchestrators and Prometheus can poll without consuming quota.
The built-in limiter is in-memory per worker. When you set workers > 1, serve() automatically divides the configured rate_limit across workers (rate_limit // workers, floor 1) so the aggregate rate stays close to the value you configured — and prints a warnings.warn at startup explaining exactly what it did.Example: serve(workers=4, config={"rate_limit": 100}) → each worker enforces 25/min locally, so the OS-load-balanced aggregate is ≈ 100/min.For exact cross-process enforcement, front the service with an API gateway or a Redis-backed limiter — the built-in in-memory limiter cannot share state across worker processes.

Recipe Serve

Programmatic server configuration and deployment

Endpoints Code

Client-side code for calling recipe endpoints