Quick Start
Configuration
Scope
Cache scope partitions collection names; it does not verify tenant ownership. Avoid sharing cached private responses across users, and include the application context in your cache design.
How It Works
- Before calling the LLM, the input is embedded and searched against the vector store
- If a result exceeds the
similarityThreshold, it’s returned as a cache hit - Output guardrails still run on cached responses
- After an LLM call, the input + output are stored in the vector store (fire-and-forget)
- TTL is enforced on lookup — expired entries are evicted lazily
Events
Supported Backends
AnyVectorStore implementation works: InMemoryVectorStore, QdrantVectorStore, MongoDBVectorStore, PgVectorStore.
Backend Examples
InMemory (Development)
Qdrant
PgVector (PostgreSQL)
Cache Hit vs Miss Behavior
This function observes an existing Agent configured with semantic caching. It supplies separate session IDs so conversation history does not become the reason for a different answer. A similar input is only a candidate hit: inspect the emitted event rather than assuming the embedding score clears the threshold.Tuning similarityThreshold
Start with
0.92 and adjust based on your cache hit rate and quality.
Cross-References
- Tool Caching — Cache individual tool results (different from semantic cache)
- Cost Tracking — Semantic cache reduces LLM costs; track savings with CostTracker