Your AI agent calls tools, returns structured data, and loops until it gets a result — but wiring all of that together by hand is a nightmare of duct-tape code. Pydantic Agent fixes that. It gives you a clean, typed interface where tools, prompts, output validation, and the agent loop itself are all first-class citizens.
In this tutorial, you will learn to:
- Create your first Pydantic Agent with system prompts and a tool
- Run agents synchronously and asynchronously
- Stream responses and listen to agent lifecycle events
- Validate structured output using Pydantic models
- Handle retries and debug agent runs with type-safe code
Page Contents
What Is Pydantic Agent?
Pydantic Agent is the primary interface in the Pydantic AI framework for building AI-powered agents in Python. Think of it as a typed, production-ready wrapper around your LLM that handles the messy parts: tool calling, message history, retries, structured output, and observability. If you are new to Pydantic, start with our Pydantic v2 guide first to understand the base model and validation system that Pydantic Agent builds on top of.
At its core, an Agent holds five things:
- Instructions — system prompts that tell the model what to do
- Tools — functions the model can call to fetch real-world data or take actions
- Output type — a Pydantic model the model must produce at the end of a run
- Dependencies — typed context (like a database connection or API client) injected at runtime
- Model — the LLM to use (OpenAI, Anthropic, Gemini, etc.)
The result is an agent that is type-safe end-to-end: your IDE catches dependency mismatches, output type errors, and tool signature mistakes before the code ever runs.
Prerequisites
You need Python 3.10+ and the pydantic-ai package:
pip install pydantic-ai
You also need an API key for your chosen model provider. Pydantic Agent supports OpenAI, Anthropic, Google Gemini, Ollama, and more via a unified model interface.
Your First Agent: Roulette Wheel
The canonical first example from the official docs is a roulette wheel checker. The winning number is passed as a dependency, and the agent uses a tool to check whether a given square is the winner:
from pydantic_ai import Agent, RunContext
roulette_agent = Agent(
'openai:gpt-4o',
deps_type=int,
output_type=bool,
system_prompt=(
'Use the `check_square` tool to determine if a given square '
'is the winning number. Ask the user which square they want to check.'
),
)
@roulette_agent.tool
async def check_square(ctx: RunContext[int], square: int) -> str:
"""Check if the square is the winning number."""
return 'winner' if square == ctx.deps else 'loser'
# The deps (winning number) is injected at run time
success_number = 18
result = roulette_agent.run_sync('I bet on square eighteen', deps=success_number)
print(result.output) #> True
result = roulette_agent.run_sync('I bet on square five', deps=success_number)
print(result.output) #> False
Key things to notice here:
- The agent is typed as
Agent[int, bool]— deps areint, output isbool - The winning number lives in
deps, not inside the tool — this keeps secrets and context separate from tool logic @roulette_agent.tooldecorates a function the LLM can callRunContext[int]gives the tool access to the deps (the winning number)result.outputis guaranteed to be abool— Pydantic validates it
Running Agents: Five Ways
Pydantic Agent gives you five ways to run an agent, depending on whether you need sync, async, or streaming:
from pydantic_ai import Agent
agent = Agent('openai:gpt-4o')
# 1. Synchronous — blocks until complete
result_sync = agent.run_sync('What is the capital of Italy?')
print(result_sync.output)
# 2. Asynchronous — use in async code
import asyncio
async def main():
result = await agent.run('What is the capital of France?')
print(result.output)
asyncio.run(main())
# 3. Async streaming — stream text as it arrives
async def stream_example():
async with agent.run_stream('What is the capital of the UK?') as response:
async for text in response.stream_text():
print(text, end='', flush=True)
asyncio.run(stream_example())
# 4. Stream all events — see tool calls + text as they happen
async def stream_events_example():
from pydantic_ai import AgentStreamEvent
async with agent.run_stream_events('What is the capital of Mexico?') as stream:
async for event in stream:
print(type(event).__name__, event)
asyncio.run(stream_events_example())
# 5. Iter — step through the agent graph node by node
async def iter_example():
async with agent.iter('What is the capital of Germany?') as agent_run:
async for node in agent_run:
print(type(node).__name__)
asyncio.run(iter_example())
For simple scripts, run_sync() is the easiest. For production async applications (FastAPI, for example), use agent.run() with async/await. For real-time UX, use streaming so the user sees output as it arrives.
Structured Output with Pydantic Models
One of Pydantic Agent’s most powerful features is forcing the model to return a specific data structure. Instead of parsing raw text, you define a Pydantic model and the agent guarantees the output matches it:
from pydantic import BaseModel
from pydantic_ai import Agent
class WeatherResult(BaseModel):
location: str
forecast: str
temperature: int
weather_agent = Agent(
'openai:gpt-4o',
output_type=WeatherResult,
system_prompt='Provide weather forecasts based on the location the user asks about.',
)
result = weather_agent.run_sync('What is the weather in Tokyo?')
print(result.output.location) #> Tokyo
print(result.output.temperature) #> e.g., 22
print(result.output.forecast) #> e.g., Sunny
The model is asked to produce JSON that matches WeatherResult. Pydantic validates it at runtime — if the model returns malformed data, Pydantic Agent retries automatically up to the configured retry limit.
Dynamic System Prompts with RunContext
Static system prompts are just strings. Dynamic prompts are functions that run at call time, so they can use runtime context — like the dependencies passed to the agent:
from datetime import date
from pydantic_ai import Agent, RunContext
class UserContext:
name: str
tier: str # 'free' or 'premium'
agent = Agent(
'openai:gpt-4o',
deps_type=UserContext,
system_prompt='You are a helpful assistant for our SaaS platform.',
)
@agent.system_prompt
def add_user_details(ctx: RunContext[UserContext]) -> str:
return f"The user's name is {ctx.deps.name} and they are on the {ctx.deps.tier} tier."
@agent.system_prompt
def add_date() -> str:
return f'Today is {date.today()}.'
result = agent.run_sync(
'What features should I upgrade to?',
deps=UserContext(name='Alice', tier='free'),
)
print(result.output)
Both dynamic prompts are appended to the system prompt at run time. The add_user_details function has access to the deps, so it can personalize the prompt with the user’s name and subscription tier.
Tool Retries and Self-Correction
Pydantic Agent handles retries in two scenarios: when a tool raises ModelRetry, and when output validation fails. This lets the model self-correct without you writing retry loops manually:
from pydantic import BaseModel
from pydantic_ai import Agent, RunContext, ModelRetry
class ChatResult(BaseModel):
user_id: int
message: str
agent = Agent(
'openai:gpt-4o',
deps_type=dict,
output_type=ChatResult,
)
@agent.tool(retries=2)
def get_user_id(ctx: RunContext[dict], name: str) -> int:
"""Look up a user ID by name."""
users = ctx.deps # a dict simulating a database
if name not in users:
raise ModelRetry(f"No user found with name {name!r}. Provide the full name.")
return users[name]
result = agent.run_sync(
'Send a message to John Doe asking for coffee',
deps={'John Doe': 123, 'Jane Smith': 456},
)
print(result.output) #> user_id=123 message='Hello John, would you be free for coffee?'
The tool retries twice (three total attempts) before the agent gives up and raises an error. You can access ctx.retry inside the tool to know which attempt you are on. The agent-level default retry count is 1, but each tool can override it.
Multi-Run Conversations
A single agent run can span an entire conversation. But for stateful conversations across separate API calls, pass the message history from one run into the next:
agent = Agent('openai:gpt-4o')
# First run — the model learns about Einstein
result1 = agent.run_sync('Who was Albert Einstein?')
print(result1.output)
#> Albert Einstein was a German-born theoretical physicist.
# Second run — pass the previous messages so the model knows who "his" refers to
result2 = agent.run_sync(
'What was his most famous equation?',
message_history=result1.new_messages(),
)
print(result2.output)
#> Albert Einstein's most famous equation is E = mc².
result1.new_messages() returns the conversation history from the first run. Pass it to the second run_sync() call and the model can reference entities from earlier messages.
Usage Limits — Prevent Infinite Loops
When running agents in production, you want to cap token usage and prevent runaway tool loops. Pydantic Agent provides UsageLimits for this:
from pydantic_ai import Agent, UsageLimits
agent = Agent('anthropic:claude-sonnet-4-20250514')
# Limit output to 10 tokens — useful for short answers
result = agent.run_sync(
'What is the capital of Italy? Answer with just the city.',
usage_limits=UsageLimits(response_tokens_limit=10),
)
print(result.output) #> Rome
# Prevent infinite tool loops — stop after 3 requests
result = agent.run_sync(
'Keep calling the tool',
usage_limits=UsageLimits(request_limit=3),
)
response_tokens_limit caps the output tokens. request_limit caps how many LLM request rounds (user message → model → response) are allowed in a single run. tool_calls_limit caps individual tool invocations.
Debugging Agent Runs
Agent runs produce detailed traces when something goes wrong. Use capture_run_messages() to inspect every message exchanged during a run:
from pydantic_ai import Agent, ModelRetry, UnexpectedModelBehavior, capture_run_messages
agent = Agent('openai:gpt-4o')
@agent.tool_plain
def calc_volume(size: int) -> int:
if size == 42:
return size ** 3
raise ModelRetry('Please try again.')
with capture_run_messages() as messages:
try:
result = agent.run_sync('Get the volume of a box with size 6.')
except UnexpectedModelBehavior as e:
print(f'Error: {e}')
print(f'Cause: {e.__cause__}')
print(f'Messages exchanged: {messages}')
When combined with Pydantic Logfire, you get a visual trace of every tool call, token usage, latency, and retry inside the web UI — invaluable for debugging production agents.
Type Safety in Practice
Pydantic Agent is designed to work with static type checkers like mypy and pyright. If you mix up the dependency type or pass the wrong output type, your IDE catches it. For a deeper grounding in Python type hints and annotations, which Pydantic Agent extends into the agentic AI space, check out our comprehensive guide.
from pydantic_ai import Agent, RunContext
# Agent expects deps=str
agent = Agent('openai:gpt-4o', deps_type=str, output_type=bool)
# Wrong — system prompt function expects RunContext[str], not RunContext[int]
@agent.system_prompt
def bad_prompt(ctx: RunContext[int]) -> str: # type error!
return f"User is {ctx.deps}"
# Wrong — output is bool, not bytes
result = agent.run_sync('Question?', deps='Frank')
result.output + b'data' # type error!
Running mypy on code like this immediately flags both errors before you ever execute the script.
Common Mistakes / Gotchas
- Confusing
system_promptwithinstructions— When you passmessage_historyto a run,instructionsare NOT retained from previous runs. Useinstructionsfor single-run isolation andsystem_promptwhen you want prompts to persist across runs. - Forgetting to escape braces in f-strings with code blocks — If you generate article content using f-strings and include example code with f-expressions like
{variable}, you must escape them as{{variable}}otherwise Python raisesNameErrorat runtime. - Not setting
deps_typewhen you need context — If your tools need runtime context (a DB connection, an API key), you must setdeps_type=YourDepsClasson the Agent and passdeps=YourDepsClass(...)at run time. Without it,RunContext.depsisNone. - Assuming
result.outputis always a string —result.outputis typed to whateveroutput_typeyou set. If you setoutput_type=bool, it is abool, not astr. Always use Pydantic models for structured outputs.
Summary & Next Steps
Pydantic Agent brings the full power of Pydantic’s type system to AI agent development. You get typed tools, typed dependencies, typed outputs, automatic retries, conversation history, streaming, and observability — all in one clean API. It integrates with Pydantic Logfire for production tracing and works with every major LLM provider. For context on how Python agents fit into the broader AI ecosystem, see our Python for AI Agents guide.
Where to go next:
- Build agents with multiple tools in a toolset
- Use
capabilitiesto bundle reusable agent behaviors - Define agents declaratively in YAML with Agent Specs
- Set up Logfire tracing for production observability
- Explore Pydantic Evals for systematic agent testing
Frequently Asked Questions
Can Pydantic Agent use models other than OpenAI?
Yes. Pydantic Agent supports OpenAI, Anthropic, Google Gemini, Azure, Bedrock, Ollama, Groq, and any provider that implements the model interface. Use the format 'provider:model-name' — for example 'anthropic:claude-sonnet-4-20250514' or 'google-gla:gemini-2.0-flash'.
How does Pydantic Agent handle retries?
Tool retries (when a tool raises ModelRetry) are tracked per-tool. Output validation retries are tracked globally. The default is 1 retry for both, configurable at the agent level, per-tool level, or per-run level via tool_retries, output_retries, and UsageLimits.
Can I run multiple agents concurrently?
Yes. Use max_concurrency on the Agent constructor to limit concurrent runs. For example, Agent('openai:gpt-4o', max_concurrency=10) allows 10 simultaneous runs and queues the rest. You can also use asyncio.gather to run multiple agent instances in parallel.

