Skip to main content
The goal loop runs your agent with its tools and only stops when an independent judge confirms the goal is met โ€” no more โ€œdoneโ€-keyword false positives.

Quick Start

1

Simple goal

Pass a task and a goal โ€” the judge decides when the goal is really met.
2

With acceptance criteria

Add structured criteria so the judge anchors to concrete evidence.
3

Resume a paused run

A paused run (max_turns exhausted) resumes automatically for the same goal โ€” turns_used is not reset.

How It Works

The loop iterates with tools; after each iteration an independent judge reads the goal plus the latest output and returns done or continue. The judge is strict about evidence:
Do NOT accept mere claims of completion (e.g. the word โ€˜doneโ€™) as evidence. Require concrete evidence that the goal is actually met.

Acceptance Criteria

GoalCriteria describes what โ€œdoneโ€ means, how to check it, and what must never happen.

run_goal() Options


Completion Reasons

The loop returns an AutonomyResult whose completion_reason tells you why it stopped.

The Independent Judge

The judge is deliberately simple, cheap, and safe to fail.
Uses a separate judge model so the acting model never self-grades โ€” this avoids self-enhancement bias.
A broken or timed-out judge yields continue, so a weak judge never wedges progress.
Only the goal plus the last ~4000 characters of agent output are judged, not the whole transcript.
Any constraint listed in GoalCriteria.constraints blocks a done verdict.
3 consecutive unparseable judge responses auto-pause the loop with budget_paused, so a wedged judge cannot silently burn the whole budget.

Async Variant

run_goal_async() mirrors run_goal() for async code.
The completion judgeโ€™s blocking litellm.completion(...) call is offloaded to a worker thread when awaited from run_goal_async() / run_autonomous_async, so it does not stall the shared event loop. Chatbots, gateways, and API servers typically share one event loop across many concurrent agents; before this change every autonomous iteration blocked that loop for a full LLM round-trip (hundreds of ms up to several seconds under rate-limiting or retries), stalling every other agentโ€™s await points for that window.The gate is cancellation-safe: cancelling the awaiting task lets the in-flight gate finish its state mutations before the cancellation propagates, so the goal state and journal never end up half-written across a teardown or a subsequent run of the same agent. Added in PR #5136.

run_until(goal=...) Delegation

Passing goal= to run_until() now delegates to run_goal() โ€” the acceptance-criteria judge gates a real tool-iteration loop instead of re-generating a whole answer.
When goal is passed, run_until() returns an AutonomyResult (success / output / completion_reason) instead of the usual EvaluationLoopResult. Adjust downstream code that accesses result.final_score / result.iterations.

Best Practices

Set judge_model to a different model than the actor so the run is not self-graded.
Use concrete, checkable evidence (โ€œexit code 0โ€, โ€œPR URL existsโ€), not a feeling.
Start with max_turns=10โ€“20 and let resume=True continue a paused run when needed.
constraints is the only field that can block a done verdict โ€” list what must never happen.

Autonomy โ€” the autonomous loop and its completion signals
Goal Engineering โ€” specify and score a goal (the Goal Loop runs and gates toward one)
Evaluation Loop โ€” the run_until Ralph Loop
SDK Reference โ€” the underlying run_autonomous loop