Skip to main content
Define operating limits before increasing traffic. One user request may produce several model calls, tool executions, and external requests.

Set limits at each boundary

Start with cost auto-stop, rate limiting, execution policy, and HTTP streaming. A limit enforced around one model call is not automatically a budget for a team or an external service.

Collect evidence you can act on

Track run status, tool outcomes, latency, token/cost accounting, cancellation, and correlated request/run IDs. Keep sensitive content out of routine telemetry unless you have explicitly configured bounded capture at the tracer and destination. Use observability for exporters and ownership. Independently configured webhooks can transmit payloads; telemetry settings do not redact them automatically.

Shut down in ownership order

  1. Stop accepting new work and stop new job claims.
  2. Drain work within a deadline; cancel connection-owned requests when their owners leave.
  3. Wait for or reconcile in-flight external actions according to the task contract.
  4. Close owned Agents, sessions, providers, browsers, and stores after their users finish.
  5. Shut down observers/exporters so queued telemetry can drain.
An application that shares a database client or store decides when to close it. A component must not unexpectedly close a borrowed service still used by another component.

Rehearse startup and drain

  1. Start the owned-session API and wait for its listener message before sending requests. For a worker, wait for your queue connection and executor registration before reporting readiness.
  2. Make one request or queue job. Retain its terminal result and any correlated trace. A listening port alone does not prove model availability.
  3. Stop admission, signal shutdown, and verify that already admitted work settles before its Agent and backing services close. Apply a host-owned deadline when dependencies can stall.
  4. Confirm that the process exits, sockets close, and telemetry artifacts are flushed. The quality-gate project demonstrates closing the Agent before shutting down its observer.
Add dependency health to readiness according to the service contract. Do not call a paid model on every liveness check.

Distinguish three kinds of failure

Execution failed: a provider error, tool error, canceled run, or stopping limit. Investigate the terminal status and the failing boundary. Execution succeeded but the answer is wrong: evaluate retrieval, tool selection, argument correctness, and answer grounding. Use evaluations, not exception counts alone. The external outcome is uncertain: preserve the operation and reconcile it. Use the recovery guide before retrying.