ReliabilityEval checks run completion, non-empty text, and expected successful tool calls. Use it when the choice of action matters in addition to the final wording.
Quick Start
This host helper receives an Agent already configured with the application’slookup_order tool. It makes model calls through that Agent; the host owns its lifecycle. Use a deterministic provider and fake service when testing the pipeline itself.
result.failed and each case’s scores. Merely mentioning a tool name in text does not pass the tool assertion. Calls with an error or denial do not count as successful expected calls.
Case Options
A tool can fail and be handled inside the Agent loop without the invocation throwing. Do not use
shouldError: true as a general assertion about a tool error. Timeouts and cancellation are evaluation failures rather than successful expected errors.
Tool Call Match Scorer
Use the scorer inside anEvalSuite when combining tool assertions with application output checks: