Skip to main content
Use TokenRateLimiter to admit estimated usage and ConcurrencyLimiter to bound active work. These are host utilities: v4 does not have an AgentConfig.rateLimit option and does not automatically switch to a cheaper model when capacity is exhausted.

Quick Start

This factory wraps an existing Agent. The caller supplies a conservative token estimate; the completed run reconciles it with reported usage. The host owns the Agent and closes it during shutdown.
Authenticate the identity before calling this function. Choose estimates for the whole Agent run, which can include multiple model calls. This is an admission control example, not a guarantee against provider charges above a limit.

Token Rate Limiter

The implementation uses fixed minute/hour windows that start with the scope’s first bucket. check(scope, estimate?) reads availability without reserving; acquire(estimate, scope) increments token and request accounting when allowed. record(actual, estimate, scope) adjusts the token reservation. It does not undo the counted request. Always supply the intended identity fields. Missing fields are omitted from the scope key. Buckets are process-local: multiple workers do not share a quota. A distributed deployment needs a shared host-side admission mechanism.

Concurrency Limiter

new ConcurrencyLimiter(maxConcurrent, timeoutMs) queues callers when all slots are occupied. acquire() resolves to a release function or rejects after its wait timeout. Always release in finally. active, pending, and available expose current process-local counts. The timeout bounds waiting for a slot; it does not cancel the operation after admission. Use the Agent’s cancellation signal and a host deadline separately.

Limit Reached Strategies

The host decides whether to return a retry response, schedule a job, or select another configured model. Although the shared RateLimitConfig type includes degradation options, TokenRateLimiter does not execute a model switch, queue model calls, or emit automatic degradation events. Test rejection with a deliberately low limit, then confirm the next request can acquire a released concurrency slot after an error. Combine this with budgets, queues, and deployment checks.