Quick Start
1
Simple Usage
Update the target model at any time β the next invocation reads the live
llm value.2
With Configuration
Clear the runtime cache after credential rotation or register a custom resolver:
How It Works
The subsystem reads the agentβs currentllm (or model) attribute at invocation time, not at construction time.
Handoffs always execute via
agent.chat() / agent.achat() directly β they do not call into this subsystem. Runtime resolution is used for agent-level model configuration and caching, not for handoff execution.Configuration Options
SessionContext
Passed to resolve_runtime to scope caching and track depth.
Cache constants
Cache keys use the format
"{session_id}:{agent_id}:{model_ref}" β caches are session-isolated so different conversations never share runtimes.
Background cleanup thread (PR #5219)
The cleanup daemon runs every_cleanup_interval seconds to evict expired entries, but it is not started at module import. As of PR #5219 (fixes #5162) it starts lazily on the first cache write β inside resolve_runtime, under _runtime_cache_lock β so two concurrent first-writers cannot each spawn a redundant daemon.
Fork safety.
import praisonaiagents and touching runtime.SessionContext no longer spawn the cleanup thread β nothing runs until the first resolve_runtime cache write. This makes the SDK safe to import at the parent of a pre-fork server (gunicorn --preload, multiprocessing, Celery prefork) as long as no resolution happens before the fork. See Python Import Safety & Fork Behaviour.Common Patterns
Mid-conversation model swap
Force cache refresh
Introspect resolved runtimes
Best Practices
Change llm before starting a new turn
Change llm before starting a new turn
Model re-resolution happens at each invocation boundary. Changing
agent.llm is effective immediately for the next call β no restart needed.Use clear_runtime_cache after credential rotation
Use clear_runtime_cache after credential rotation
The 5-minute TTL means old runtimes may linger after you rotate API keys. Call
clear_runtime_cache() to evict all entries and force fresh connections.Implement supports_model narrowly in custom resolvers
Implement supports_model narrowly in custom resolvers
Return
False from supports_model for models you do not handle. The built-in DefaultRuntimeResolver acts as the final fallback, so returning False simply delegates back to it.Handoffs execute via agent.chat() β not this subsystem
Handoffs execute via agent.chat() β not this subsystem
If you need to change how a handoff target executes, configure the target agent itself (instructions, tools, llm). The target agentβs full
chat() pipeline runs on every handoff β this subsystem is not in that path.Public API
All names are importable frompraisonaiagents.runtime:
resolve_runtime
RuntimeProtocol / AgentRuntimeProtocol
Protocols that custom runtimes must satisfy. Only relevant when building a custom resolver.
Related
Runtime Selection
Choose which runtime executes each model
Agent Handoffs
Agent-to-agent delegation β handoffs always use agent.chat()
Import Safety (Python)
Why the cleanup thread starts lazily and how it enables fork-safe imports

