Skip to main content
Use google(modelId, config?) from @agentium/core for Gemini generation through @google/genai. For native realtime audio, use the separate Google Live adapter.

Setup

Factory

Save as gemini-agent.ts and run with npx tsx gemini-agent.ts.

Supported models

The current general example is gemini-3.8-flash, checked against Google’s model catalog on October 4, 2026. See Example models and the Gemini catalog for model selection. A preview model should be labeled as such rather than presented as a stable default.

Config

Pass apiKey to override GOOGLE_API_KEY. Use the separate Vertex adapter for Google Cloud project/location authentication; regional model availability is a separate deployment choice.

Reasoning

Gemini 3 uses a thinking level derived from reasoning.effort; older Gemini 2.5 examples use a token budget. Agentium removes temperature and topP for Gemini 3.8 requests. Do not translate an older token budget into an assumed equivalent effort. Tool conversations preserve provider thought signatures and function-call identity in providerExtras. Keep these when persisting history and do not replay foreign provider continuation.

Multi-modal support

Use multimodal content for text, image, audio, and supported file parts. The model and selected endpoint determine valid MIME types and size limits; arbitrary spreadsheet or archive formats are not automatically understood because they are attached as a file.

Files & documents

Load only files your application has authorized, specify the actual MIME type, and bound uploads. Model input validation is separate from filesystem authorization.

Realtime / voice (Gemini Live)

GoogleLiveProvider uses a native Live session, with 16 kHz input and 24 kHz output audio. google("gemini-3.8-flash") is the text-model path and does not create an audio session.

Full example

Use a Gemini ModelProvider in the research harness, or combine the text provider with a recognizer and synthesizer in StreamingVoicePipeline.