KnowledgeBase. The knowledge base delegates vector embedding to that store; it does not infer an embedding model from a dimension number.
Use HashEmbedding for deterministic local wiring checks and a measured semantic model for real retrieval. Model availability and media support remain provider contracts.
EmbeddingProvider interface
embedBatch() returns a vector per text input. embedMultimodal() returns one vector for its combined input.
Try a local embedder
Use an ESM TypeScript project ("type": "module" in package.json) on a supported Node version.
embedding.ts and run npx tsx embedding.ts:
OpenAI embeddings
Installopenai, set OPENAI_API_KEY, and run this with an account that can use the selected embedding model:
Available models
OpenAI documentstext-embedding-3-small and text-embedding-3-large, with default vector sizes 1,536 and 3,072 respectively. The adapter also retains a dimension mapping for legacy text-embedding-ada-002. Provider guidance was checked on 2026-10-04. OpenAI embedding guide.
The adapter accepts apiKey, model, and dimensions. Its defaults are OPENAI_API_KEY, text-embedding-3-small, and that model’s known dimension. If you override dimensions, make the vector backend match and re-index existing documents as necessary.
Google embeddings
Install@google/genai and set GOOGLE_API_KEY. This example explicitly selects gemini-embedding-2; the v4 adapter’s constructor default is still text-embedding-004 when model is omitted.
Available models
Google’s current guide describesgemini-embedding-2 for text and multimodal input and gemini-embedding-001 for text. Both default to 3,072 dimensions. Checked 2026-10-04. Google embedding guide.
The adapter accepts apiKey, model, and dimensions. Its dimension lookup also includes older text-embedding-004 and embedding-001 IDs at 768; that lookup is not a guarantee those models remain available to an account. The adapter’s supportsMultimodal flag is based on a model-name prefix, so also verify the selected model’s actual provider capabilities.
Multimodal embeddings (Gemini Embedding 2)
This complete example indexes an image with a caption and searches with another image. It requiresphoto.jpg and query.jpg in the working directory, Google credentials, and the optional client above. It makes live embedding requests and keeps its index only in memory.
Index an image with a caption
Search by image
PassContentPart[] as the search query, as shown above. A plain numeric vector bypasses query embedding; a string uses text embedding. The stored document and query must use the same embedding space and dimensions.
Supported modalities
Core’s multimodal inputs usetext, image, audio, and file parts. For a supported Google embedding model, the adapter converts image/audio/video/PDF parts into the provider request. partsFromFile(path, mimeType?) reads a file and returns one ContentPart; fetchAsBase64(url) fetches a remote resource and returns data and MIME type. Select trusted source files/URLs and respect the provider’s current size and media limits.
Output dimensions (Matryoshka)
Google documents reduced dimensions for its embedding models. Configure the requested dimension explicitly and keep the store consistent:Important: v1 and v2 vectors are NOT interchangeable
Changing the model changes the embedding space, even when vector lengths match. Re-index documents when switching model/version, dimensions, or incompatible input preparation. Keep the prior index until the replacement has passed your retrieval checks.Limitations
- Multimodal embedding returns one vector for the combined input; call separately when each item needs its own vector.
- The vector adapters retain text, vectors, and metadata, not original media parts. Store useful source references in metadata.
- Provider acceptance, account permissions, media limits, and live latency are not established by TypeScript compilation.
- Batch behavior is adapter-specific; the Google wrapper issues per-input requests in bounded groups rather than a managed provider batch job.
Using with KnowledgeBase
This complete local program uses the same provider → vector store → knowledge-base composition as a live setup:KnowledgeBase.close() closes its store, so coordinate ownership if multiple consumers share it. Continue with the retrieval tutorial to pass evidence into an Agent, or vector stores to choose persistence.