New in 2026: Master Python for AI, Data Science

Python

Pydantic Agent Basics: A Complete 2026 Tutorial

Pydantic Agent Basics: A Complete 2026 Tutorial

Your AI agent calls tools, returns structured data, and loops until it gets a result — but wiring all of that together by hand is a nightmare of duct-tape code. Pydantic Agent fixes that. It gives you a clean, typed interface where tools, prompts, output validation, and the agent loop itself are all first-class citizens.

In this tutorial, you will learn to:

  • Create your first Pydantic Agent with system prompts and a tool
  • Run agents synchronously and asynchronously
  • Stream responses and listen to agent lifecycle events
  • Validate structured output using Pydantic models
  • Handle retries and debug agent runs with type-safe code

What Is Pydantic Agent?

Pydantic Agent is the primary interface in the Pydantic AI framework for building AI-powered agents in Python. Think of it as a typed, production-ready wrapper around your LLM that handles the messy parts: tool calling, message history, retries, structured output, and observability. If you are new to Pydantic, start with our Pydantic v2 guide first to understand the base model and validation system that Pydantic Agent builds on top of.

At its core, an Agent holds five things:

  • Instructions — system prompts that tell the model what to do
  • Tools — functions the model can call to fetch real-world data or take actions
  • Output type — a Pydantic model the model must produce at the end of a run
  • Dependencies — typed context (like a database connection or API client) injected at runtime
  • Model — the LLM to use (OpenAI, Anthropic, Gemini, etc.)

The result is an agent that is type-safe end-to-end: your IDE catches dependency mismatches, output type errors, and tool signature mistakes before the code ever runs.

Prerequisites

You need Python 3.10+ and the pydantic-ai package:

pip install pydantic-ai

You also need an API key for your chosen model provider. Pydantic Agent supports OpenAI, Anthropic, Google Gemini, Ollama, and more via a unified model interface.

Your First Agent: Roulette Wheel

The canonical first example from the official docs is a roulette wheel checker. The winning number is passed as a dependency, and the agent uses a tool to check whether a given square is the winner:

from pydantic_ai import Agent, RunContext

roulette_agent = Agent(
    'openai:gpt-4o',
    deps_type=int,
    output_type=bool,
    system_prompt=(
        'Use the `check_square` tool to determine if a given square '
        'is the winning number. Ask the user which square they want to check.'
    ),
)


@roulette_agent.tool
async def check_square(ctx: RunContext[int], square: int) -> str:
    """Check if the square is the winning number."""
    return 'winner' if square == ctx.deps else 'loser'


# The deps (winning number) is injected at run time
success_number = 18
result = roulette_agent.run_sync('I bet on square eighteen', deps=success_number)
print(result.output)  #> True

result = roulette_agent.run_sync('I bet on square five', deps=success_number)
print(result.output)  #> False

Key things to notice here:

  • The agent is typed as Agent[int, bool] — deps are int, output is bool
  • The winning number lives in deps, not inside the tool — this keeps secrets and context separate from tool logic
  • @roulette_agent.tool decorates a function the LLM can call
  • RunContext[int] gives the tool access to the deps (the winning number)
  • result.output is guaranteed to be a bool — Pydantic validates it

Running Agents: Five Ways

Pydantic Agent gives you five ways to run an agent, depending on whether you need sync, async, or streaming:

from pydantic_ai import Agent

agent = Agent('openai:gpt-4o')

# 1. Synchronous — blocks until complete
result_sync = agent.run_sync('What is the capital of Italy?')
print(result_sync.output)

# 2. Asynchronous — use in async code
import asyncio

async def main():
    result = await agent.run('What is the capital of France?')
    print(result.output)

asyncio.run(main())

# 3. Async streaming — stream text as it arrives
async def stream_example():
    async with agent.run_stream('What is the capital of the UK?') as response:
        async for text in response.stream_text():
            print(text, end='', flush=True)

asyncio.run(stream_example())

# 4. Stream all events — see tool calls + text as they happen
async def stream_events_example():
    from pydantic_ai import AgentStreamEvent
    async with agent.run_stream_events('What is the capital of Mexico?') as stream:
        async for event in stream:
            print(type(event).__name__, event)

asyncio.run(stream_events_example())

# 5. Iter — step through the agent graph node by node
async def iter_example():
    async with agent.iter('What is the capital of Germany?') as agent_run:
        async for node in agent_run:
            print(type(node).__name__)

asyncio.run(iter_example())

For simple scripts, run_sync() is the easiest. For production async applications (FastAPI, for example), use agent.run() with async/await. For real-time UX, use streaming so the user sees output as it arrives.

Structured Output with Pydantic Models

One of Pydantic Agent’s most powerful features is forcing the model to return a specific data structure. Instead of parsing raw text, you define a Pydantic model and the agent guarantees the output matches it:

from pydantic import BaseModel
from pydantic_ai import Agent

class WeatherResult(BaseModel):
    location: str
    forecast: str
    temperature: int

weather_agent = Agent(
    'openai:gpt-4o',
    output_type=WeatherResult,
    system_prompt='Provide weather forecasts based on the location the user asks about.',
)

result = weather_agent.run_sync('What is the weather in Tokyo?')
print(result.output.location)   #> Tokyo
print(result.output.temperature)  #> e.g., 22
print(result.output.forecast)   #> e.g., Sunny

The model is asked to produce JSON that matches WeatherResult. Pydantic validates it at runtime — if the model returns malformed data, Pydantic Agent retries automatically up to the configured retry limit.

Dynamic System Prompts with RunContext

Static system prompts are just strings. Dynamic prompts are functions that run at call time, so they can use runtime context — like the dependencies passed to the agent:

from datetime import date
from pydantic_ai import Agent, RunContext

class UserContext:
    name: str
    tier: str  # 'free' or 'premium'

agent = Agent(
    'openai:gpt-4o',
    deps_type=UserContext,
    system_prompt='You are a helpful assistant for our SaaS platform.',
)


@agent.system_prompt
def add_user_details(ctx: RunContext[UserContext]) -> str:
    return f"The user's name is {ctx.deps.name} and they are on the {ctx.deps.tier} tier."


@agent.system_prompt
def add_date() -> str:
    return f'Today is {date.today()}.'


result = agent.run_sync(
    'What features should I upgrade to?',
    deps=UserContext(name='Alice', tier='free'),
)
print(result.output)

Both dynamic prompts are appended to the system prompt at run time. The add_user_details function has access to the deps, so it can personalize the prompt with the user’s name and subscription tier.

Tool Retries and Self-Correction

Pydantic Agent handles retries in two scenarios: when a tool raises ModelRetry, and when output validation fails. This lets the model self-correct without you writing retry loops manually:

from pydantic import BaseModel
from pydantic_ai import Agent, RunContext, ModelRetry

class ChatResult(BaseModel):
    user_id: int
    message: str

agent = Agent(
    'openai:gpt-4o',
    deps_type=dict,
    output_type=ChatResult,
)


@agent.tool(retries=2)
def get_user_id(ctx: RunContext[dict], name: str) -> int:
    """Look up a user ID by name."""
    users = ctx.deps  # a dict simulating a database
    if name not in users:
        raise ModelRetry(f"No user found with name {name!r}. Provide the full name.")
    return users[name]


result = agent.run_sync(
    'Send a message to John Doe asking for coffee',
    deps={'John Doe': 123, 'Jane Smith': 456},
)
print(result.output)  #> user_id=123 message='Hello John, would you be free for coffee?'

The tool retries twice (three total attempts) before the agent gives up and raises an error. You can access ctx.retry inside the tool to know which attempt you are on. The agent-level default retry count is 1, but each tool can override it.

Multi-Run Conversations

A single agent run can span an entire conversation. But for stateful conversations across separate API calls, pass the message history from one run into the next:

agent = Agent('openai:gpt-4o')

# First run — the model learns about Einstein
result1 = agent.run_sync('Who was Albert Einstein?')
print(result1.output)
#> Albert Einstein was a German-born theoretical physicist.

# Second run — pass the previous messages so the model knows who "his" refers to
result2 = agent.run_sync(
    'What was his most famous equation?',
    message_history=result1.new_messages(),
)
print(result2.output)
#> Albert Einstein's most famous equation is E = mc².

result1.new_messages() returns the conversation history from the first run. Pass it to the second run_sync() call and the model can reference entities from earlier messages.

Usage Limits — Prevent Infinite Loops

When running agents in production, you want to cap token usage and prevent runaway tool loops. Pydantic Agent provides UsageLimits for this:

from pydantic_ai import Agent, UsageLimits

agent = Agent('anthropic:claude-sonnet-4-20250514')

# Limit output to 10 tokens — useful for short answers
result = agent.run_sync(
    'What is the capital of Italy? Answer with just the city.',
    usage_limits=UsageLimits(response_tokens_limit=10),
)
print(result.output)  #> Rome

# Prevent infinite tool loops — stop after 3 requests
result = agent.run_sync(
    'Keep calling the tool',
    usage_limits=UsageLimits(request_limit=3),
)

response_tokens_limit caps the output tokens. request_limit caps how many LLM request rounds (user message → model → response) are allowed in a single run. tool_calls_limit caps individual tool invocations.

Debugging Agent Runs

Agent runs produce detailed traces when something goes wrong. Use capture_run_messages() to inspect every message exchanged during a run:

from pydantic_ai import Agent, ModelRetry, UnexpectedModelBehavior, capture_run_messages

agent = Agent('openai:gpt-4o')


@agent.tool_plain
def calc_volume(size: int) -> int:
    if size == 42:
        return size ** 3
    raise ModelRetry('Please try again.')


with capture_run_messages() as messages:
    try:
        result = agent.run_sync('Get the volume of a box with size 6.')
    except UnexpectedModelBehavior as e:
        print(f'Error: {e}')
        print(f'Cause: {e.__cause__}')
        print(f'Messages exchanged: {messages}')

When combined with Pydantic Logfire, you get a visual trace of every tool call, token usage, latency, and retry inside the web UI — invaluable for debugging production agents.

Type Safety in Practice

Pydantic Agent is designed to work with static type checkers like mypy and pyright. If you mix up the dependency type or pass the wrong output type, your IDE catches it. For a deeper grounding in Python type hints and annotations, which Pydantic Agent extends into the agentic AI space, check out our comprehensive guide.

from pydantic_ai import Agent, RunContext

# Agent expects deps=str
agent = Agent('openai:gpt-4o', deps_type=str, output_type=bool)

# Wrong — system prompt function expects RunContext[str], not RunContext[int]
@agent.system_prompt
def bad_prompt(ctx: RunContext[int]) -> str:  # type error!
    return f"User is {ctx.deps}"

# Wrong — output is bool, not bytes
result = agent.run_sync('Question?', deps='Frank')
result.output + b'data'  # type error!

Running mypy on code like this immediately flags both errors before you ever execute the script.

Common Mistakes / Gotchas

  • Confusing system_prompt with instructions — When you pass message_history to a run, instructions are NOT retained from previous runs. Use instructions for single-run isolation and system_prompt when you want prompts to persist across runs.
  • Forgetting to escape braces in f-strings with code blocks — If you generate article content using f-strings and include example code with f-expressions like {variable}, you must escape them as {{variable}} otherwise Python raises NameError at runtime.
  • Not setting deps_type when you need context — If your tools need runtime context (a DB connection, an API key), you must set deps_type=YourDepsClass on the Agent and pass deps=YourDepsClass(...) at run time. Without it, RunContext.deps is None.
  • Assuming result.output is always a string — result.output is typed to whatever output_type you set. If you set output_type=bool, it is a bool, not a str. Always use Pydantic models for structured outputs.

Summary & Next Steps

Pydantic Agent brings the full power of Pydantic’s type system to AI agent development. You get typed tools, typed dependencies, typed outputs, automatic retries, conversation history, streaming, and observability — all in one clean API. It integrates with Pydantic Logfire for production tracing and works with every major LLM provider. For context on how Python agents fit into the broader AI ecosystem, see our Python for AI Agents guide.

Where to go next:

  • Build agents with multiple tools in a toolset
  • Use capabilities to bundle reusable agent behaviors
  • Define agents declaratively in YAML with Agent Specs
  • Set up Logfire tracing for production observability
  • Explore Pydantic Evals for systematic agent testing

Frequently Asked Questions

Can Pydantic Agent use models other than OpenAI?

Yes. Pydantic Agent supports OpenAI, Anthropic, Google Gemini, Azure, Bedrock, Ollama, Groq, and any provider that implements the model interface. Use the format 'provider:model-name' — for example 'anthropic:claude-sonnet-4-20250514' or 'google-gla:gemini-2.0-flash'.

How does Pydantic Agent handle retries?

Tool retries (when a tool raises ModelRetry) are tracked per-tool. Output validation retries are tracked globally. The default is 1 retry for both, configurable at the agent level, per-tool level, or per-run level via tool_retries, output_retries, and UsageLimits.

Can I run multiple agents concurrently?

Yes. Use max_concurrency on the Agent constructor to limit concurrent runs. For example, Agent('openai:gpt-4o', max_concurrency=10) allows 10 simultaneous runs and queues the rest. You can also use asyncio.gather to run multiple agent instances in parallel.

Related posts
ProgrammingPython

Production-Ready MCP Servers — Security, Testing & Deployment

ProgrammingPython

Build Your First MCP Server with Python SDK — Fundamentals

ProgrammingPython

Connect FastAPI to MCP — Two Integration Patterns

Python

ValueError: Length of values does not match length of index — Python Error Causes and How to Fix It

Leave a Reply