Skip to main content
Build a ticket classifier and check that invoices route to billing while delivery questions route to shipping. This recipe combines core, eval, and observability into a development loop you can use before shipping an application. The fixture tests the wiring, scorer, artifacts, and failure handling. A separate live evaluation tests actual model behavior.

Get the project

Download the complete project, extract it, and run npm install. Then run npm run check and npm start. For manual setup in an empty directory:

Create the application

The schema defines the contract, the case’s expected value defines the answer, and the scorer checks each case independently. The observer records execution metadata. Each invocation writes to a new artifacts/quality-* directory so previous runs remain available. Save as quality-gate.ts:
quality-gate.ts

Verify a pass and a failure

The default run prints passed: 2/2; failed: 0 and exits successfully. The --broken fixture deliberately routes the delivery case incorrectly: expect passed: 1/2; failed: 1 and exit code 1. This proves the process fails on a quality regression even when the Agent run itself completed. Inspect the printed directory:
  • report.json contains case names, scorer decisions, errors, counts, and metrics.
  • traces.jsonl contains one compact trace per completed run, with model/run correlation and metadata. Content capture is off by default.

Connect a live model

Only --live selects OpenAI. Model access and outputs can vary; see model choices. Add representative cases and keep the scorer tied to your application’s output contract. To change providers, pass a different ModelProvider to runQualityGate().

Use the result in CI

The downloaded project’s npm start preserves the process exit code. Run it as a normal CI step after npm ci or npm install. Upload the artifacts directory even when the gate fails. Keep the offline fixture check and any credentialed evaluation as separate jobs so their evidence stays clear. Use tool-call scoring for action selection, conversation tests for multi-turn behavior, and an LLM judge for assertions that require a model. A text match is not sufficient evidence for those tasks.

Troubleshoot and adapt

After active runs settle, the program closes the Agent and shuts down telemetry to drain queued exports. Keep useful reports, then remove generated artifact directories when you no longer need them. Next, serve the application or follow Ship.