Skip to main content
An evaluation scores your agent’s output so you can compare prompts or models and catch regressions. It has three parts:
  • Dataset: test cases, with inputs and expected outputs.
  • Task: the change you want to test, such as a new prompt.
  • Evaluator: code or an LLM judge that scores each output.
The example below runs two experiments over three airline-policy questions. The first prompt scores 0; the second adds the policy and scores 1. You need an OpenAI API key. To have a coding agent or PXI do this instead, see Connect Your Coding Agent.

Run two experiments

Save as evaluate.py and run python evaluate.py.

Compare the runs

Open the dataset’s Experiments tab to compare the two runs side by side. A higher score on the same dataset and judge means the change helped; a run that still scores 0 has the judge’s explanation attached.

Learn more

The code on this page lives in examples/quickstarts, ready to run.