Skip to main content
Turn the running gateway into an OpenAI-compatible and MCP endpoint so SDK clients and MCP tools reach the same live agents as chat users. One process, one agent state, three protocols. This is different from praisonai serve openai, which runs a separate standalone OpenAI-only process.

Quick Start

1

Enable in gateway.yaml

Add the api block to your gateway config:
2

Start with CLI flags

Enable the same surfaces from the command line:
3

Enable in Python

Pass constructor flags to the gateway:

How It Works

Each API request dispatches into the gateway’s own registered agents, sharing the same session store and admission gate as chat users.

Endpoints Exposed

Call the OpenAI surface with any standard client:

Background Runs

Submit long agent turns in the background, retrieve the result by id, and cancel in-flight runs — over the same OpenAI-compatible surface. The synchronous path is unchanged and remains the default — omit background and store to keep the exact previous behaviour.

Which mode fits

Quick Start

1

Submit in the background

2

Poll until terminal

3

Cancel if needed

User Flow

Status Vocabulary

A completed response always carries an assistant message (even with empty text), so a client reading output[0] keeps working. Polling the same completed response returns the same output[0].id each time.

Response Object

The error block appears only when status is failed. A GET on an unknown id returns 404 with {"error": {"type": "invalid_request_error", "message": "No such response: resp-…"}}.
When the caller uses a stable identity (a bearer token, or an OpenAI-Session / X-Session-Id header), only that caller can retrieve or cancel their own background runs. A different token receives 404 (not 403, to avoid leaking id existence). Anonymous callers on auth-disabled gateways are looked up by id alone — the same open posture that applies to every /v1/* route without auth.
Background responses live in a bounded in-process store (cap: 1024 entries). When full, the oldest terminal entry (completed/cancelled/failed) is evicted first; in-flight runs are never evicted. Results do not survive a gateway restart — this is intentional for now; durable cross-restart persistence is tracked separately.
Both new routes go through the same _api_guarded wrapper as POST /v1/responses (gated by gateway.auth_token) and only mount when gateway.api.openai = true.

Token Usage

Every response reports real per-turn token counts from the agent’s LLM instance. Each surface reports usage in its own OpenAI-native shape: When an agent exposes no LLM metrics, usage falls back to zeros in the correct spec shape, so clients never need to special-case a missing block.

Streamed usage

Opt in with stream_options.include_usage: true to receive a streamed usage chunk — the same contract OpenAI’s own API uses.
The usage-only chunk is emitted right before data: [DONE]. Without include_usage, the stream frames are unchanged and no extra chunk is sent. This is orthogonal to token-level streaming below — it works with or without it.

Token-level streaming (opt-in)

Enable gateway.api.stream to emit each DELTA_TEXT / FIRST_TOKEN piece as its own chat.completion.chunk with delta.content, so time-to-first-content matches token latency instead of full-turn latency. Default is off: without opt-in the SSE surface is byte-for-byte the previous single buffered chunk, so existing clients keep working unchanged. Enable it in gateway.yaml:
Or in Python:
Consume it from any standard OpenAI client — no client-side change needed:
Behaviour guarantees:

Configuration Options

The gateway.api block maps to the ApiConfig dataclass. All three flags are opt-in; when openai and mcp are both False (default), no extra routes are mounted. ApiConfig also exposes an enabled property (true if either surface is on), plus to_dict() and from_dict().
Full API surface

How Auth Works

Every API route is protected by the same gateway.auth_token as /info and /metrics. The /info endpoint advertises which surfaces are enabled in its api field, so clients can introspect a running gateway before connecting.

When to Use vs praisonai serve openai

Use the gateway api: block when you want SDK clients and MCP tools to share live agents and sessions with chat users. Use praisonai serve openai when you want a standalone, lightweight OpenAI-only process with no gateway state. See OpenAI-Compatible Server.

Best Practices

Leave openai: false and mcp: false (the defaults) unless a client needs them. Disabled surfaces mount no routes and leave the gateway unchanged.
Every /v1/* and /mcp route uses the same gateway.auth_token. Set a strong token when binding to any non-loopback interface.
Pass an OpenAI-Session or X-Session-Id header to reuse one agent session across calls. Without it, stateless callers get a fresh session per request.
GET /info returns an api field listing enabled surfaces, so tooling can confirm what a gateway exposes before dispatching.
Pass stream_options={"include_usage": True} when using stream=True if your client tracks cost per response. The extra chunk arrives right before [DONE] and carries real per-turn totals.
The gateway snapshots the agent’s per-turn token metrics in the same execution context that produced the reply, so a usage field returned to one caller cannot be corrupted by a concurrent turn on the same shared agent. No configuration required.
Any turn that could exceed the client’s HTTP timeout is a candidate. POST /v1/responses with background: true returns a persisted id in ~milliseconds; poll GET /v1/responses/{id} at a comfortable interval (2–5s) until status is terminal. If the caller loses connectivity, the run keeps going and the result is still retrievable by id (subject to the in-process store cap).
Owner-scoped retrieval only kicks in when the caller has a stable identity. Anonymous callers on auth-disabled gateways are looked up by id alone. Set gateway.auth_token in any deployment where multiple clients share the same gateway.
/cancel fires the same interrupt seam as the WebSocket /stop path — the agent checks it at its own checkpoints and stops between LLM calls / tool calls. A tight, non-yielding loop inside a tool won’t stop until it yields; the task-cancel fallback covers most cases but not synchronous CPU-bound work.

Gateway

Gateway architecture and YAML configuration

OpenAI-Compatible Server

Standalone OpenAI-only server process

MCP Integration

Model Context Protocol servers and clients

Gateway CLI

CLI commands for managing the gateway