> ## Documentation Index
> Fetch the complete documentation index at: https://docs.agentium.in/llms.txt
> Use this file to discover all available pages before exploring further.

# Build a quality gate with traces

> Evaluate structured answers, save a report and traces, and fail CI on a deliberate regression.

Build a ticket classifier and check that invoices route to billing while delivery questions route to shipping. This recipe combines **core**, **eval**, and **observability** into a development loop you can use before shipping an application.

| You will use | Requirement |
| - | - |
| Packages | `@agentium/core`, `@agentium/eval`, `@agentium/observability`, `zod` |
| Default execution | Deterministic local model fixture; no credentials or network |
| Optional live execution | `openai`, `OPENAI_API_KEY`, a compatible text model |
| Outputs | Per-case report, JSONL traces, and a pass/fail process exit code |

The fixture tests the wiring, scorer, artifacts, and failure handling. A separate live evaluation tests actual model behavior.

## Get the project

[Download the complete project](/downloads/quality-gate.zip), extract it, and run `npm install`. Then run `npm run check` and `npm start`.

For manual setup in an empty directory:

```bash theme={null}
npm init -y
npm pkg set type=module
npm install @agentium/core@4.0.0 @agentium/eval@4.0.0 @agentium/observability@4.0.0 zod
npm install --save-dev typescript tsx @types/node
```

## Create the application

The schema defines the contract, the case's `expected` value defines the answer, and the scorer checks each case independently. The observer records execution metadata. Each invocation writes to a new `artifacts/quality-*` directory so previous runs remain available.

Save as `quality-gate.ts`:

```typescript quality-gate.ts theme={null}
import { mkdir, mkdtemp, writeFile } from "node:fs/promises";
import { join } from "node:path";
import { pathToFileURL } from "node:url";
import { Agent, openai, type ModelProvider } from "@agentium/core";
import { custom, EvalSuite } from "@agentium/eval";
import { instrument, JsonFileExporter } from "@agentium/observability";
import { z } from "zod";

const classification = z.object({ queue: z.enum(["billing", "shipping"]) });

// A deterministic provider makes the evaluation pipeline runnable offline.
function createFixture(broken: boolean): ModelProvider {
  return {
    providerId: "fixture", modelId: "ticket-classifier",
    async generate(messages) {
      const input = String(messages.filter((message) => message.role === "user").at(-1)?.content ?? "");
      const queue = !broken && input.includes("delivery") ? "shipping" : "billing";
      return {
        message: { role: "assistant", content: JSON.stringify({ queue }) },
        finishReason: "stop", raw: null,
        usage: { promptTokens: 10, completionTokens: 5, totalTokens: 15 },
      };
    },
    async *stream() { throw new Error("This recipe uses generate(), not stream()"); },
  };
}

export async function runQualityGate(model: ModelProvider, outputRoot = "artifacts") {
  await mkdir(outputRoot, { recursive: true });
  const directory = await mkdtemp(join(outputRoot, "quality-"));
  const agent = new Agent({
    name: "ticket-classifier", model, register: false,
    instructions: "Classify invoices as billing and delivery questions as shipping.",
    structuredOutput: classification,
  });
  const telemetry = instrument(agent, {
    exporters: [new JsonFileExporter({ path: join(directory, "traces.jsonl") })],
  });
  try {
    const result = await new EvalSuite({
      name: "ticket-routing", agent, concurrency: 1, timeoutMs: 10_000,
      cases: [
        { name: "invoice", input: "Explain my invoice", expected: "billing" },
        { name: "delivery", input: "Where is my delivery?", expected: "shipping" },
      ],
      scorers: [custom("correct-queue", (_input, output, expected) => {
        const parsed = classification.safeParse(output.structured);
        const pass = parsed.success && parsed.data.queue === expected;
        return { score: pass ? 1 : 0, pass, reason: pass ? "Correct queue" : `Expected ${expected}` };
      })],
    }).run();
    const summary = {
      passed: result.passed, failed: result.failed, total: result.total,
      cases: result.results.map((item) => ({ name: item.caseName, pass: item.pass, scores: item.scores, error: item.error })),
      metrics: telemetry.metrics?.getMetrics(),
    };
    // Await the write so report failures cannot silently produce a passing gate.
    await writeFile(join(directory, "report.json"), JSON.stringify(summary, null, 2));
    console.log(`passed: ${result.passed}/${result.total}; failed: ${result.failed}`);
    console.log("artifacts:", directory);
    return { failed: result.failed, directory };
  } finally {
    try { await agent.close(); } finally { await telemetry.shutdown(); }
  }
}

if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) {
  const live = process.argv.includes("--live");
  if (live && !process.env.OPENAI_API_KEY) throw new Error("Set OPENAI_API_KEY before using --live");
  const model = live
    ? openai(process.env.OPENAI_MODEL ?? "gpt-6-luna")
    : createFixture(process.argv.includes("--broken"));
  const result = await runQualityGate(model);
  if (result.failed > 0) process.exitCode = 1;
}
```

## Verify a pass and a failure

```bash theme={null}
npx tsx quality-gate.ts
npx tsx quality-gate.ts --broken
```

The default run prints `passed: 2/2; failed: 0` and exits successfully. The `--broken` fixture deliberately routes the delivery case incorrectly: expect `passed: 1/2; failed: 1` and **exit code 1**. This proves the process fails on a quality regression even when the Agent run itself completed.

Inspect the printed directory:

* `report.json` contains case names, scorer decisions, errors, counts, and metrics.
* `traces.jsonl` contains one compact trace per completed run, with model/run correlation and metadata. Content capture is off by default.

## Connect a live model

```bash theme={null}
npm install openai
export OPENAI_API_KEY="your-key"
export OPENAI_MODEL="gpt-6-luna"
npx tsx quality-gate.ts --live
```

Only `--live` selects OpenAI. Model access and outputs can vary; see [model choices](/reference/example-models). Add representative cases and keep the scorer tied to your application's output contract. To change providers, pass a different `ModelProvider` to `runQualityGate()`.

## Use the result in CI

The downloaded project's `npm start` preserves the process exit code. Run it as a normal CI step after `npm ci` or `npm install`. Upload the `artifacts` directory even when the gate fails. Keep the offline fixture check and any credentialed evaluation as separate jobs so their evidence stays clear.

Use [tool-call scoring](/eval/reliability) for action selection, [conversation tests](/eval/conversational) for multi-turn behavior, and an [LLM judge](/eval/agent-judge) for assertions that require a model. A text match is not sufficient evidence for those tasks.

## Troubleshoot and adapt

| Symptom | Check |
| - | - |
| Completed Agent run, failed evaluation | Inspect the failing case's scorer; execution success and answer correctness are different |
| Live mode fails before a case | Check credentials, model access, and the explicit `--live` flag |
| Deadline failures | Review the case timeout and provider latency before increasing concurrency |
| Empty trace content | Metadata-only capture is intentional; use bounded [capture settings](/observability/overview) if your application needs content |
| Artifact write fails | Check directory permissions; report write errors propagate instead of silently passing |

After active runs settle, the program closes the Agent and shuts down telemetry to drain queued exports. Keep useful reports, then remove generated artifact directories when you no longer need them. Next, [serve the application](/examples/authenticated-api) or follow [Ship](/ship/overview).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.