Skip to main content

Overview

AgentJudgeEval evaluates agent responses against multiple custom criteria using an LLM judge. Supports both numeric scoring (0.0–1.0) and binary (PASS/FAIL) modes. The judge must be a chat model. Do not pass jev() here — this evaluator expects a numeric or PASS/FAIL response. To score with Jev, use custom().

Quick Start

Install @agentium/core@4.0.0, @agentium/eval@4.0.0, and the optional openai SDK. Set OPENAI_API_KEY; this example makes live provider calls. For a no-key runnable project, start with the quality-gate recipe.

Scoring Modes

  • numeric (default): Each criterion scored 0.0–1.0
  • binary: Each criterion scored PASS (1.0) or FAIL (0.0)