Skip to main content
Agentium is built for minimal overhead — fewer tokens, faster responses, lower cost. This page covers the key optimizations and benchmark results.

Benchmark Results

All benchmarks use gpt-4o-mini, identical prompts, and 5 runs per scenario. Agentium and LangChain run on Node.js; Agno runs on Python.

Simple Completion

Tool Calling

Agentium and LangChain produce identical tool schemas (167 prompt tokens). Agentium strips verbose JSON Schema metadata ($schema, additionalProperties) to keep schemas compact.

Multi-turn Memory

Agentium uses 39% fewer prompt tokens and 43% less cost than LangChain for multi-turn conversations. LangChain injects heavier system prompts and history formatting overhead.

Summary

Agentium is the fastest for tool calling, the cheapest for multi-turn conversations, and matches LangChain on tool schema efficiency. Response latency is within noise across simple completions.

Optimizations

Tool Schema Caching & Optimization

Tool definitions (Zod-to-JSON Schema conversion) are computed once at agent construction and cached. Verbose JSON Schema metadata ($schema, additionalProperties, description on the root object) is stripped automatically — reducing token overhead without losing semantic information.
For OpenAI models, tools can opt into strict mode for guaranteed valid JSON output:

Automatic Retry

Transient LLM API failures are automatically retried with exponential backoff + jitter. Retryable errors include HTTP 429, 5xx, and network errors.
Default: 3 retries, 500ms initial delay, 10s max delay.

Token-Based History Trimming

Trim history on memory, not on Agent:
There is no maxContextTokens on Agent.

Background extraction

When memory.userFacts (or profile / entities / learnings) is on, extraction runs after the answer is sent. The user is not waiting on it. Tiny always-on notes (fileMemory) are injected into the prompt instead of being fetched as a tool.

Streaming Usage Tracking

Token usage (promptTokens, completionTokens, totalTokens, reasoningTokens) is accurately tracked in both run() and stream() modes. Stream usage is accumulated from provider finish chunks.

Methodology

  • All benchmarks use gpt-4o-mini with identical prompts
  • Each scenario runs 5 times; results are averaged
  • Startup time measures framework import + agent initialization
  • Cost uses gpt-4o-mini pricing: 0.15/1Minput,0.15/1M input, 0.6/1M output
  • Network latency to OpenAI is shared across all frameworks
  • Full benchmark scripts are in benchmarks/ in the repository