Skip to main content
Semantic caching stores LLM responses indexed by the semantic meaning of the input. When a similar query arrives, the cached response is returned without calling the LLM — reducing costs and latency.

Quick Start

Configuration

Scope

Cache scope partitions collection names; it does not verify tenant ownership. Avoid sharing cached private responses across users, and include the application context in your cache design.

How It Works

  1. Before calling the LLM, the input is embedded and searched against the vector store
  2. If a result exceeds the similarityThreshold, it’s returned as a cache hit
  3. Output guardrails still run on cached responses
  4. After an LLM call, the input + output are stored in the vector store (fire-and-forget)
  5. TTL is enforced on lookup — expired entries are evicted lazily

Events

Supported Backends

Any VectorStore implementation works: InMemoryVectorStore, QdrantVectorStore, MongoDBVectorStore, PgVectorStore.

Backend Examples

InMemory (Development)

The in-memory cache is lost when the process restarts. Use it for a local experiment with a bounded set of public queries.

Qdrant

PgVector (PostgreSQL)


Cache Hit vs Miss Behavior

This function observes an existing Agent configured with semantic caching. It supplies separate session IDs so conversation history does not become the reason for a different answer. A similar input is only a candidate hit: inspect the emitted event rather than assuming the embedding score clears the threshold.
Cache writes happen asynchronously. The second request can miss if the first write has not settled. Use cache events and representative paraphrases to measure hits. A cache hit still requires an embedding lookup and can become stale when the underlying data or prompt changes.

Tuning similarityThreshold

Start with 0.92 and adjust based on your cache hit rate and quality.

Cross-References

  • Tool Caching — Cache individual tool results (different from semantic cache)
  • Cost Tracking — Semantic cache reduces LLM costs; track savings with CostTracker