Quick Start
Configuration
boolean
required
Enable or disable reasoning for this agent.
'none' | 'minimal' | 'low' | 'medium' | 'high' | 'xhigh' | 'max'
How hard the model thinks. OpenAI-family and Gemini 3 use it directly. Claude 4.6+ and Claude 5 map it to
output_config.effort (minimal becomes low). DeepSeek only accepts low / high / max, so medium is sent as high. Mistral sends high or none. "none" turns thinking off where the provider allows it. Grok 4.5 and 4.6 ignore none.number
Token budget. Used by Claude models that still take
thinking.budget_tokens (before Opus/Sonnet 4.6), Gemini 2.5, and Cohere tokenBudget. Default on those Claude models: 10000.'auto' | 'concise' | 'detailed'
OpenAI Responses summary. Default
"detailed", which is what fills result.thinking. "auto" asks for the most detailed summary that model supports.'standard' | 'pro'
GPT-5.6 and GPT-6 Responses execution mode.
pro spends more work and latency.'auto' | 'current_turn' | 'all_turns'
Which OpenAI reasoning items later turns should see. Omit it to keep the model default.
Provider Support
Each provider mapsReasoningConfig to its native API:
When reasoning is enabled for OpenAI,
temperature is ignored (the API does not allow both). For Anthropic, temperature and top_p are stripped when thinking is active.Tools + reasoning
GPT-5.4+, GPT-5.6, and GPT-6 reject function tools on Chat Completions unlessreasoning_effort is "none". GPT-5.6 defaults to medium when you omit the field, which is why an agent with tools 400s even if you never set reasoning.
Agentium routes that combination to /v1/responses and keeps your effort. If the endpoint is missing (old Azure API version, some proxies), it falls back to Chat Completions with reasoning_effort: "none" so the run still works.
providerExtras) and sends it back. You do not pass it yourself.
OpenAI only returns readable thinking when the request sets reasoning.summary. Agentium sends "detailed" unless you set summary.
Output
When reasoning is enabled,RunOutput includes:
Streaming
During streaming, reasoning content is delivered asthinking chunks:
ReasoningConfig Type
providerOptions is separate from reasoning:
promptCacheputs an Anthropiccache_controlbreakpoint on the system prompt.compactionTokensturns on Anthropic server compaction (compact_20260112). Values under 50000 are raised to 50000.clearToolResultsturns on Anthropic tool-result clearing.promptCacheRetentionis OpenAI Responsesprompt_cache_retention.mediaResolution,cachedContent, andgoogleSearchare Gemini generate options. Gemini 3.8 also dropstemperatureandtopP.
Logging
WhenlogLevel is "info" or higher, the agent logger displays:
- Thinking content (truncated, italic, dimmed)
- Reasoning token count with a brain icon in the usage summary
Accessing Thinking Content
When reasoning is enabled, the model’s internal thinking is available inRunOutput.thinking:
thinking content is never shown to end users by default — it’s available for debugging, logging, or advanced use cases.
Reasoning with maxTokens
On Claude models that still usebudget_tokens (before Opus/Sonnet 4.6), maxTokens must accommodate both thinking and response tokens. Agentium handles this automatically:
- If
maxTokensis set but too small for the thinking budget, it’s overridden tobudgetTokens + 4096 - For example:
maxTokens: 1024withbudgetTokens: 2000results in an effectivemax_tokensof6096sent to the Anthropic API
max_tokens, set maxTokens to a value larger than budgetTokens + 1024.
Logging Reasoning Output
EnablelogLevel: "debug" to see thinking content in the console:
"info" level, only token usage summaries are logged. At "debug", the full thinking content is printed.