Skip to main content
ReliabilityEval checks run completion, non-empty text, and expected successful tool calls. Use it when the choice of action matters in addition to the final wording.

Quick Start

This host helper receives an Agent already configured with the application’s lookup_order tool. It makes model calls through that Agent; the host owns its lifecycle. Use a deterministic provider and fake service when testing the pipeline itself.
Inspect result.failed and each case’s scores. Merely mentioning a tool name in text does not pass the tool assertion. Calls with an error or denial do not count as successful expected calls.

Case Options

A tool can fail and be handled inside the Agent loop without the invocation throwing. Do not use shouldError: true as a general assertion about a tool error. Timeouts and cancellation are evaluation failures rather than successful expected errors.

Tool Call Match Scorer

Use the scorer inside an EvalSuite when combining tool assertions with application output checks:
This checks invocation evidence, not that an external system applied an effect. Verify service receipts and idempotency in your application tests. Continue with conversation trajectories when tool order across multiple turns matters.