Skip to main content
@agentium/eval runs an Agent against named cases and scores each result. Start with an application contract: a correct queue, a required tool call, an allowed answer, or a latency limit. A completed run is not automatically a correct answer. For an offline project with pass/fail commands, reports, and traces, use the quality-gate recipe.

Quick Start

Install the packages and the optional provider SDK. Set OPENAI_API_KEY before running this live-model example with tsx in an ESM TypeScript project.
Save as eval.ts. Each case supplies its own expected answer; the scorer reads that case’s expectation.
If both answers match, expect 2/2 passed and exit code 0. A wrong answer, failed run, or case deadline produces a failure. Exact text is appropriate for this constrained task; use structured output or semantic criteria when valid answers can vary.

Built-in Scorers

Every scorer in a suite runs on every case. contains("Paris") always checks for Paris; it does not automatically read EvalCase.expected. Use custom() to consume per-case expectations. Scorer names must be unique. Keep failure reasons useful for a developer reading a CI report. A substring match cannot prove a payment occurred or an authorization check was enforced; test those application effects directly.

LLM-as-Judge

Use a chat model capable of the judge’s required response format. A judge adds provider calls and may disagree across runs. Calibrate it against reviewed cases. AgentJudgeEval accepts your own criterion strings; AccuracyEval compares answers with expected values.

Jev as a judge

Jev is a decision model. Use the Jev custom-scorer guide when its typed decision questions match your rubric; do not substitute it blindly into a chat judge.

Reporters

ConsoleReporter prints results; new JsonReporter("./eval-results.json") writes a JSON report. The result contains name, results, passed, failed, total, averageScore, and durationMs. Each case has pass, scores, timing, and optional error/output. There are no allPassed, failures, or passRate properties. JsonReporter logs write failures rather than throwing. If CI must fail when artifacts cannot be written, await your own writeFile() as the quality-gate recipe does. Inspect per-case failureKind and cleanupPending when diagnosing cancellation or deadlines; a timeout cannot undo an external effect.

Custom Scorer

Start with observable application data: validate output.structured, inspect successful tool calls, or compare a deterministic field with expected. Return a finite score from 0 to 1 and an explicit pass decision. Keep network or database checks cancellable and bounded by the evaluation deadline.

Full Eval Suite Example

The complete quality-gate project demonstrates structured classification, a deterministic fixture, a deliberate regression, explicit artifact writes, and telemetry shutdown. Extend its cases before replacing the fixture with a live provider.

Running Evals in CI

Gate on result.failed > 0 and set process.exitCode = 1 after writing artifacts. Let cleanup finish; avoid an immediate process.exit() that can cut off exports. Separate offline pipeline checks from credentialed evaluations and record the model, package version, cases, and rubric used. Next, test tool reliability, multi-turn conversations, or performance.