knowledge.search(...) and task-callback memory writes are offloaded automatically — you don’t call anything new. What you see is parallel async tasks that share a knowledge base or memory now actually running in parallel.
Quick Start
1
Parallel agents sharing knowledge
An
asyncio.gather(...) of agents that query the same knowledge base runs in parallel — no lookup blocks the loop.2
Parallel tasks writing memory
Task-callback memory writes offload to worker threads, so a fan-out of tasks doesn’t stall on one slow write.
async_execution=True only fans tasks out under the async entry point shown here (astart() / arun()). The sync .start() / .run() runs them sequentially and logs a warning — see PraisonAI PR #5227.3
Parallel agents with guardrails
An
asyncio.gather(...) of agents where one carries a string guardrail runs in parallel — the guardrail’s validation LLM call offloads to a worker thread instead of stalling the fan-out.How It Works
What runs off the event loop
Both were synchronous embedding + vector-store / DB calls that previously blocked the event loop for the whole coroutine lifetime. An
asyncio.gather(...) fan-out would serialise on the slowest lookup or write instead of running concurrently. Offloading them to threads keeps the loop free.
A string / LLMGuardrail output guardrail fires a blocking LLM call under the hood to validate the agent’s output. On the async path this call is now offloaded via loop.run_in_executor(...), so parallel tasks don’t serialise on one task’s guardrail round-trip (PR #4469). Contextvars are preserved across the offload with copy_context_to_callable, so trace emission and session context stay intact.
Session auto-save (Agent(memory=MemoryConfig(auto_save="…"))) does a FileLock + JSON read-modify-write on the process-wide DefaultSessionStore singleton. Before PR #5258 that synchronous write ran inline on the async call sites, so one async agent’s auto-save blocked every other async agent sharing the loop — a fast agent gathered alongside a slow-writing one would visibly stall. The write is now offloaded via asyncio.to_thread at every async call site, and an internal lock serialises the shared _auto_save_last_index bookkeeping so two concurrent async runs sharing one Agent can’t double-persist the same turns.
This is transparent — there is no user-facing knob to opt out. You do not call anything new; parallel async tasks simply run in parallel.
Thread-safe in-memory adapter
The built-inInMemoryAdapter guards its store, search, delete, and reset methods with a re-entrant lock (RLock). Once writes move off the event loop, parallel async writes from an asyncio.gather(...) fan-out cannot produce duplicate ids or lose entries.
If you write a custom adapter that will be shared across async agents, do the same — guard its mutating methods with a lock.
See also (PR #5219): two more concurrency guarantees landed in PR #5219 (fixes #5162):
- Image Generation → Non-blocking async —
ImageAgent.achat/agenerate_imagenow offload viaasyncio.to_thread, so agatherof an image agent and a text agent no longer starves the text agent. - Thread Safety → Output singleton lock —
output="status"/output="trace"metrics no longer corrupt underasyncio.gather.
See also (PR #5370): Thread-Safe Agent State → AsyncSafeState reentrancy fix — the sync
with state: entry point now uses DualLock.sync(), closing a same-thread self-deadlock on nested chat-history mutations.Best Practices
Fan out with asyncio.gather without fear
Fan out with asyncio.gather without fear
Knowledge search and memory writes offload to threads, so a
gather(...) of agents that share a knowledge base or memory runs in parallel instead of serialising on the slowest task.Make custom adapters thread-safe
Make custom adapters thread-safe
Custom memory adapters used from async task callbacks receive writes from multiple worker threads. Guard mutating methods with a lock, as the built-in
InMemoryAdapter does with an RLock.Guardrails don't stall async fan-outs
Guardrails don't stall async fan-outs
String /
LLMGuardrail output and task guardrails offload their blocking LLM call to a worker thread on the async path, so an asyncio.gather(...) of agents that share a guardrail no longer serialises on one task’s validation LLM round-trip. Custom guardrail objects that implement synchronous validate_output get the same offload for free.async def guardrails work on the sync path too
async def guardrails work on the sync path too
Since PraisonAI PR #5305, passing an
async def guardrail to Agent(guardrails=...) or Task(guardrails=...) works on both the sync and async execution paths. _process_guardrail detects a coroutine function and runs it via run_coroutine_from_any_context. Before #5305 the sync path called the coroutine function synchronously — the returned coroutine was never awaited, GuardrailResult.from_tuple crashed on it, and the run was reported as a validation failure. Retries were wasted on a guardrail that had never actually executed.Nothing to configure
Nothing to configure
The offload is automatic and transparent. There is no flag to enable or disable it — the sync path is unchanged and the async path always stays responsive.
Related
Knowledge
Add document knowledge with async-safe search
Memory
Persistent memory with offloaded, thread-safe writes
Async Agents
Run agents and tasks concurrently with asyncio
Custom Memory Adapters
Build your own adapter — guard it with a lock for async use

