The fixture tests the wiring, scorer, artifacts, and failure handling. A separate live evaluation tests actual model behavior.
Get the project
Download the complete project, extract it, and runnpm install. Then run npm run check and npm start.
For manual setup in an empty directory:
Create the application
The schema defines the contract, the case’sexpected value defines the answer, and the scorer checks each case independently. The observer records execution metadata. Each invocation writes to a new artifacts/quality-* directory so previous runs remain available.
Save as quality-gate.ts:
quality-gate.ts
Verify a pass and a failure
passed: 2/2; failed: 0 and exits successfully. The --broken fixture deliberately routes the delivery case incorrectly: expect passed: 1/2; failed: 1 and exit code 1. This proves the process fails on a quality regression even when the Agent run itself completed.
Inspect the printed directory:
report.jsoncontains case names, scorer decisions, errors, counts, and metrics.traces.jsonlcontains one compact trace per completed run, with model/run correlation and metadata. Content capture is off by default.
Connect a live model
--live selects OpenAI. Model access and outputs can vary; see model choices. Add representative cases and keep the scorer tied to your application’s output contract. To change providers, pass a different ModelProvider to runQualityGate().
Use the result in CI
The downloaded project’snpm start preserves the process exit code. Run it as a normal CI step after npm ci or npm install. Upload the artifacts directory even when the gate fails. Keep the offline fixture check and any credentialed evaluation as separate jobs so their evidence stays clear.
Use tool-call scoring for action selection, conversation tests for multi-turn behavior, and an LLM judge for assertions that require a model. A text match is not sufficient evidence for those tasks.
Troubleshoot and adapt
After active runs settle, the program closes the Agent and shuts down telemetry to drain queued exports. Keep useful reports, then remove generated artifact directories when you no longer need them. Next, serve the application or follow Ship.