How It Works
Quick Start
1
Simple Usage
Enable rate limiting and metrics on an existing server:Scrape metrics at
GET /metrics (Prometheus format).2
With Configuration
Production setup with auth, admin reload, workers, and tracing:
How It Works
Advanced features layer onto the base recipe server from Recipe Serve. Rate limiting uses a sliding window per client; metrics expose request counts and latency; admin endpoints hot-reload recipes without restart.Rate limiter internals
The middleware buckets requests by the strongest identity the auth layer has already validated — never by a raw client header — so rotating a header cannot defeat the limit.RateLimiter exposes two entry points:
check(client_id)— synchronous, thread-safe (backward-compatible).check_syncremains as an alias.check_async(client_id)— async, guarded by anasyncio.Lock. Used by the middleware hot path so overlapping requests can never both pass the length check before either appends.
If you rely on rate limiting in
auth: none or auth: jwt mode, the bucket id is derived from the validated identity (JWT sub or client IP), never from the raw X-API-Key header. A client cannot rotate X-API-Key values to escape the limit. In auth: api-key mode the validated key is the bucket, so distinct valid keys naturally get distinct buckets — this is intentional.jwt:<sub> > apikey:<key> (only in api-key mode) > ip:<host> > "anonymous".
Configuration Options
Common Patterns
Programmatic rate limiter
- Sync
- Async
Admin reload
Prometheus scrape config
OpenTelemetry dependencies
Best Practices
Always enable auth with admin endpoints
Always enable auth with admin endpoints
Set
enable_admin=True only alongside auth: api-key and load the key from PRAISONAI_API_KEY. Admin reload can change live behaviour — protect it.Use workers for CPU-bound recipe loads
Use workers for CPU-bound recipe loads
Set
workers to roughly 2 × CPU cores + 1. Workers above 1 disable hot reload automatically. Under the hood, serve(workers=N>1) hands uvicorn an app-factory import string (praisonai.recipe.serve:_app_factory) and ships the config to each worker through the PRAISONAI_RECIPE_SERVE_CONFIG environment variable. Previously, passing workers=N silently degraded to a single worker — that regression is fixed.Exempt health and metrics from rate limits
Exempt health and metrics from rate limits
Keep
/health and /metrics in rate_limit_exempt_paths so orchestrators and Prometheus can poll without consuming quota.Use Redis for distributed rate limiting
Use Redis for distributed rate limiting
The built-in limiter is in-memory per worker. When you set
workers > 1, serve() automatically divides the configured rate_limit across workers (rate_limit // workers, floor 1) so the aggregate rate stays close to the value you configured — and prints a warnings.warn at startup explaining exactly what it did.Example: serve(workers=4, config={"rate_limit": 100}) → each worker enforces 25/min locally, so the OS-load-balanced aggregate is ≈ 100/min.For exact cross-process enforcement, front the service with an API gateway or a Redis-backed limiter — the built-in in-memory limiter cannot share state across worker processes.Related
Recipe Serve
Programmatic server configuration and deployment
Endpoints Code
Client-side code for calling recipe endpoints

