Examples¶
The repository ships eight small, self-contained projects under examples/: one per feature. Each is a known-good starting point to copy from: CI exercises every one of them (see tests/test_examples.py), so they never drift from the current release.
Every example except voice-livekit uses a deterministic mock agent, so it runs offline with no API key. Swap the llm_eval_agent fixture for a real adapter to point any of them at your own agent.
Running an example¶
cd examples/single-turn
pip install pytest-agent-eval # or: uv add pytest-agent-eval
pytest --agent-eval-live
Two CI caveats
- Judge rubrics are stubbed in CI. Examples that use an LLM-as-judge (
multi-turn-judge,tool-call-args) need an API key to run the real judge standalone. voice-livekitis collect-only in CI. It needs live credentials and synthesized audio, so CI only verifies that it collects.
The sections below follow the same pedagogical order as the examples README: start at the top and each one adds a single concept.
single-turn¶
The minimal setup: one YAML transcript plus one llm_eval_agent fixture. Reach for this shape first: a single user turn with deterministic substring and tool-call checks is enough for most smoke tests, and it needs no Python test file at all.
multi-turn-judge¶
A multi-turn conversation where the second turn is graded by an LLM-as-judge rubric. Reach for this when correctness depends on context carried across turns: here the assistant must reschedule the existing booking without asking the user to repeat themselves, which no substring check can verify.
id: booking_reschedule
threshold: 1.0
runs: 1
turns:
- user: "Book me a table for two tomorrow at 10am."
expect:
reply_contains_any: [booked, confirmed]
tool_calls_include: [create_booking]
- user: "Actually, can you move it to 11am instead?"
expect:
tool_calls_include: [update_booking]
tool_calls_exclude: [create_booking]
judge:
rubric: >
The reply confirms the new 11am time AND references the original
10am booking, without asking the user to repeat information.
import pytest
from pytest_agent_eval import AgentReply
@pytest.fixture
def llm_eval_agent():
# Deterministic mock agent that remembers nothing but answers plausibly per turn.
async def agent(history):
message = history[-1].content.lower()
if "11am" in message:
return AgentReply(
"Done, moved your booking from 10am to 11am. Reference stays BK-1234.", ["update_booking"]
)
return AgentReply("Booked for tomorrow at 10am. Reference BK-1234.", ["create_booking"])
return agent
Runnable project ↗ · See Evaluators for the judge rubric surface.
tool-calls¶
Assertions on which tools the agent invoked: include, exclude, and order. Reach for this when the reply text is not enough and you need to verify the agent took the right actions: authenticate before fetching, never call a destructive tool, and do it all in sequence.
Runnable project ↗ · See Evaluators for ToolCallEvaluator.
tool-call-args¶
Goes one level deeper than tool-calls: it asserts on the arguments the agent passed. Reach for this when calling the right tool is not enough: the party size, date, and time have to be correct too. The example shows all three modes side by side: subset (default), exact, and an LLM-judged rubric over the call's JSON arguments. Note that the fixture returns ToolCall(name, args) objects rather than plain strings, and that is what makes the arguments available to assert on.
id: booking_argument_checks
threshold: 1.0
runs: 1
turns:
- user: "Book me a table for two tomorrow at 10am."
expect:
tool_calls_include: [create_booking]
tool_calls_args:
# subset (default): these keys/values must appear; extras are fine
- tool: create_booking
args:
time: 10am
party_size: 2
# exact: the full argument dict must match
- tool: create_booking
mode: exact
args:
time: 10am
date: tomorrow
party_size: 2
# LLM-judged: rubric evaluated against the call's JSON arguments
- tool: create_booking
judge:
rubric: "The booking time must be within business hours (9am-5pm)."
import pytest
from pytest_agent_eval import AgentReply, ToolCall
@pytest.fixture
def llm_eval_agent():
# Return ToolCall(name, args) instead of plain strings to enable argument assertions.
async def agent(history):
call = ToolCall("create_booking", {"time": "10am", "date": "tomorrow", "party_size": 2})
return AgentReply("Booked for two, tomorrow at 10am!", [call])
return agent
Runnable project ↗ · See Evaluators for the full tool_calls_args specification.
regex-contains¶
Deterministic reply assertions: substring all_of plus regex matching. Reach for this when the reply must contain structured tokens (a reference number in a known format, a time, an order ID) that you can pin down with a pattern instead of paying for a judge.
Runnable project ↗ · See Evaluators for ContainsEvaluator.
python-parametrize¶
The Python API instead of YAML, combined with @pytest.mark.parametrize. Reach for this when your eval cases are data-driven: one transcript shape run across many inputs (cities, locales, product IDs), where hand-writing a YAML file per case would be repetitive. You get the full expressiveness of pytest: fixtures, marks, and parametrization.
The agent_eval argument is injected by the plugin's fixture; annotating it as EvalSession (from pytest_agent_eval.runner) makes that explicit and unlocks editor autocompletion for .run(...). Because from __future__ import annotations defers annotation evaluation, the import lives under TYPE_CHECKING: it is only needed by type checkers and editors, never at runtime.
from __future__ import annotations
from typing import TYPE_CHECKING
import pytest
from pytest_agent_eval import AgentReply, Expect, Turn
if TYPE_CHECKING:
from pytest_agent_eval.runner import EvalSession
@pytest.mark.parametrize("city", ["Ghent", "Leuven"])
@pytest.mark.agent_eval(threshold=1.0, runs=1)
async def test_booking_mentions_city(agent_eval: EvalSession, city: str) -> None:
# Deterministic mock agent; replace with your real agent or an adapter.
async def agent(history):
return AgentReply(f"Booked a table in {city}! Reference BK-1234.", ["create_booking"])
result = await agent_eval.run(
agent=agent,
turns=[
Turn(
user=f"Book me a table in {city} tomorrow at 10am.",
expect=Expect(
reply_contains_all=[city],
tool_calls_include=["create_booking"],
),
)
],
)
result.assert_threshold()
Runnable project ↗ · See the Python API for Turn and Expect.
groups¶
Group-level pass thresholds with a per-group exit-code override. Reach for this when a whole suite should be judged as a set rather than test-by-test: allow a fraction of edge cases to fail while still gating the merge on a critical case that must always pass. Groups are configured in pyproject.toml and tag their transcripts.
import pytest
from pytest_agent_eval import AgentReply
@pytest.fixture
def llm_eval_agent():
# Handles the happy path, "fails" on the deliberately hard edge case.
async def agent(history):
message = history[-1].content.lower()
if "gluten-free" in message:
return AgentReply("I'm not sure I can help with that.", [])
return AgentReply("Booking confirmed! Reference BK-1234.", ["create_booking"])
return agent
The booking group requires 50% of its gate:booking transcripts to pass, so the deliberately-failing edge_case is absorbed, but must_pass still forces booking_happy_path to pass regardless of the threshold.
Runnable project ↗ · See Group thresholds for the full configuration surface.
voice-livekit¶
Voice evals: each turn declares an audio: WAV that the LiveKitAdapter streams into a fresh LiveKit AgentSession; the tool calls and assistant transcript feed the same evaluators as a text agent. Reach for this when you are testing a real-time voice agent and want the same threshold-based scoring you use for text.
import pytest
from pytest_agent_eval.adapters.livekit import LiveKitAdapter
def make_session():
# Imported inside the factory so collection works without live credentials.
from livekit.agents.voice import Agent, AgentSession
from livekit.plugins import openai
session = AgentSession(llm=openai.realtime.RealtimeModel())
agent = Agent(instructions="You are a friendly booking assistant.", tools=[])
return session, agent
@pytest.fixture
def llm_eval_agent():
return LiveKitAdapter(make_session)
The WAV fixtures are generated from each turn's user text (hash-cached, idempotent) before running the suite:
Runnable project ↗ · See the LiveKit adapter docs for full options.