Benchmark Results
All benchmarks usegpt-4o-mini, identical prompts, and 5 runs per scenario. Agentium and LangChain run on Node.js; Agno runs on Python.
Simple Completion
Tool Calling
Agentium and LangChain produce identical tool schemas (167 prompt tokens). Agentium strips verbose JSON Schema metadata (
$schema, additionalProperties) to keep schemas compact.
Multi-turn Memory
Agentium uses 39% fewer prompt tokens and 43% less cost than LangChain for multi-turn conversations. LangChain injects heavier system prompts and history formatting overhead.
Summary
Agentium is the fastest for tool calling, the cheapest for multi-turn conversations, and matches LangChain on tool schema efficiency. Response latency is within noise across simple completions.
Optimizations
Tool Schema Caching & Optimization
Tool definitions (Zod-to-JSON Schema conversion) are computed once at agent construction and cached. Verbose JSON Schema metadata ($schema, additionalProperties, description on the root object) is stripped automatically — reducing token overhead without losing semantic information.
Automatic Retry
Transient LLM API failures are automatically retried with exponential backoff + jitter. Retryable errors include HTTP 429, 5xx, and network errors.Token-Based History Trimming
Trim history on memory, not onAgent:
maxContextTokens on Agent.
Background extraction
Whenmemory.userFacts (or profile / entities / learnings) is on, extraction runs after the answer is sent. The user is not waiting on it.
Tiny always-on notes (fileMemory) are injected into the prompt instead of being fetched as a tool.
Streaming Usage Tracking
Token usage (promptTokens, completionTokens, totalTokens, reasoningTokens) is accurately tracked in both run() and stream() modes. Stream usage is accumulated from provider finish chunks.
Methodology
- All benchmarks use
gpt-4o-miniwith identical prompts - Each scenario runs 5 times; results are averaged
- Startup time measures framework import + agent initialization
- Cost uses gpt-4o-mini pricing: 0.6/1M output
- Network latency to OpenAI is shared across all frameworks
- Full benchmark scripts are in
benchmarks/in the repository