Skip to main content
ConversationSuite runs named scenarios through an Agent and a synthetic user model. Use it for behavior that depends on earlier turns. The simulator and judge make additional model calls; this is not an offline unit test.

Quick Start

This host helper receives an Agent and a chat model for simulation. The scenario asks for two details before producing a delivery summary; it requires no external-effect tools. The host owns both provider configuration and the Agent lifecycle.
Configure the Agent instructions for that task before calling the helper. Read the recorded turns and judge decisions rather than accepting the aggregate count alone. Synthetic-user success is model-derived evidence, not proof that a real user completed a workflow.

Synthetic Users

SyntheticUser receives a persona and a model. The persona supplies a description, goal, and optional turn limit. The simulation can signal goal completion; the runner also scores the stated success criteria. Keep sensitive or production data out of test personas.

Trajectory Scoring

Add expectedTrajectory to a scenario when the Agent has real configured tools. It supports requiredTools, orderedTools, forbiddenTools, and maxToolCalls. A scenario requiring send_reset_email will fail if the supplied Agent has no such tool. Use a fake email service in evaluation; a conversational test should not unexpectedly send real messages.

Agent Comparison

ConversationRunner.runComparison(agentA, agentB, scenario) runs both Agents and returns agentA.result, agentB.result, winner, and reasoning. Results are not returned as resultA or resultB. Use equivalent provider settings and independently owned sessions, then inspect both transcripts.

Suite Results

The suite returns name, results, passed, failed, total, averageTurns, averageScore, and durationMs. Each scenario result includes turns and any trajectory match. Fail a CI gate when failed > 0, persist the transcript/report, and drain active work before closing the Agent. For a deterministic no-key starting point, use the quality-gate recipe. For single-run action assertions, see reliability.