Recipe Serve Advanced CLI
Advanced CLI options for the recipe server including rate limiting, metrics, admin endpoints, workers, and OpenTelemetry tracing.
Quick Reference
Command Options
Rate Limiting
Protect your server from abuse.
When combined with --workers N, the configured --rate_limit is divided across workers (rate_limit // workers, floor 1) so the aggregate stays close to the value you set. A startup warning prints exactly what value each worker enforces. For exact cross-process limiting, front the service with a shared-store limiter (API gateway or Redis).
Test Rate Limiting
Expected output after 5 requests:
Request Size Limits
Prevent oversized payloads.
Test Size Limit
Expected response:
Metrics Endpoint
Expose Prometheus metrics.
Sample Output
Prometheus Integration
Admin Endpoints
Hot-reload recipes without restart.
Response
Workers
Scale with multiple processes.
As of PraisonAI PR #5231, --workers N > 1 actually spawns N uvicorn worker processes; previously the CLI accepted the flag but silently ran a single worker.
Notes
- Workers > 1 automatically disables
--reload
- Each worker has independent rate limiter state;
--rate_limit is split across workers (rate_limit // workers, floor 1) with a startup warning
- For distributed rate limiting, use external store (Redis)
OpenTelemetry Tracing
Distributed tracing support.
Install Dependencies
OpenAPI Specification
Get the API specification.
Configuration File
All CLI options can be set in serve.yaml:
Use with:
Production Examples
Basic Production
Full Production
Docker
Kubernetes
Environment Variables
Troubleshooting
Rate Limit Not Working
Check if path is exempt:
Metrics Endpoint 404
Enable metrics:
Admin Endpoint 401
Provide authentication:
Workers with Reload
Cannot use both:
Next Steps