Skip to main content
Agentium supports extended thinking (reasoning) across all major model providers. When enabled, the model produces an internal chain-of-thought before generating the final answer — improving performance on complex math, logic, and multi-step problems.

Quick Start


Configuration

boolean
required
Enable or disable reasoning for this agent.
'none' | 'minimal' | 'low' | 'medium' | 'high' | 'xhigh' | 'max'
How hard the model thinks. OpenAI-family and Gemini 3 use it directly. Claude 4.6+ and Claude 5 map it to output_config.effort (minimal becomes low). DeepSeek only accepts low / high / max, so medium is sent as high. Mistral sends high or none. "none" turns thinking off where the provider allows it. Grok 4.5 and 4.6 ignore none.
number
Token budget. Used by Claude models that still take thinking.budget_tokens (before Opus/Sonnet 4.6), Gemini 2.5, and Cohere tokenBudget. Default on those Claude models: 10000.
'auto' | 'concise' | 'detailed'
OpenAI Responses summary. Default "detailed", which is what fills result.thinking. "auto" asks for the most detailed summary that model supports.
'standard' | 'pro'
GPT-5.6 and GPT-6 Responses execution mode. pro spends more work and latency.
'auto' | 'current_turn' | 'all_turns'
Which OpenAI reasoning items later turns should see. Omit it to keep the model default.

Provider Support

Each provider maps ReasoningConfig to its native API:
When reasoning is enabled for OpenAI, temperature is ignored (the API does not allow both). For Anthropic, temperature and top_p are stripped when thinking is active.

Tools + reasoning

GPT-5.4+, GPT-5.6, and GPT-6 reject function tools on Chat Completions unless reasoning_effort is "none". GPT-5.6 defaults to medium when you omit the field, which is why an agent with tools 400s even if you never set reasoning. Agentium routes that combination to /v1/responses and keeps your effort. If the endpoint is missing (old Azure API version, some proxies), it falls back to Chat Completions with reasoning_effort: "none" so the run still works.
Claude, Gemini, OpenAI Responses, and DeepSeek tool turns must echo the previous reasoning payload. Agentium stores it on the assistant message (providerExtras) and sends it back. You do not pass it yourself. OpenAI only returns readable thinking when the request sets reasoning.summary. Agentium sends "detailed" unless you set summary.

Output

When reasoning is enabled, RunOutput includes:

Streaming

During streaming, reasoning content is delivered as thinking chunks:

ReasoningConfig Type

providerOptions is separate from reasoning:
  • promptCache puts an Anthropic cache_control breakpoint on the system prompt.
  • compactionTokens turns on Anthropic server compaction (compact_20260112). Values under 50000 are raised to 50000.
  • clearToolResults turns on Anthropic tool-result clearing.
  • promptCacheRetention is OpenAI Responses prompt_cache_retention.
  • mediaResolution, cachedContent, and googleSearch are Gemini generate options. Gemini 3.8 also drops temperature and topP.

Logging

When logLevel is "info" or higher, the agent logger displays:
  • Thinking content (truncated, italic, dimmed)
  • Reasoning token count with a brain icon in the usage summary

Accessing Thinking Content

When reasoning is enabled, the model’s internal thinking is available in RunOutput.thinking:
The thinking content is never shown to end users by default — it’s available for debugging, logging, or advanced use cases.

Reasoning with maxTokens

On Claude models that still use budget_tokens (before Opus/Sonnet 4.6), maxTokens must accommodate both thinking and response tokens. Agentium handles this automatically:
  • If maxTokens is set but too small for the thinking budget, it’s overridden to budgetTokens + 4096
  • For example: maxTokens: 1024 with budgetTokens: 2000 results in an effective max_tokens of 6096 sent to the Anthropic API
If you want precise control over max_tokens, set maxTokens to a value larger than budgetTokens + 1024.

Logging Reasoning Output

Enable logLevel: "debug" to see thinking content in the console:
At "info" level, only token usage summaries are logged. At "debug", the full thinking content is printed.