Skip to main content
EdgeRuntime adds periodic resource monitoring, a health endpoint, and watchdog signals around a supplied Agent. Your application still decides when to accept work, when to heartbeat, and how to respond to a timeout. It does not restart the Agent, change its model, shed tools, or enforce a process-memory limit. Start with a working device application, then make the host’s response to these signals explicit.

Quick start

This host integration factory borrows an existing Agent. It pauses admission when the runtime reports degradation and asks a host-supplied supervisor to handle watchdog timeouts. Calling it starts monitoring and the health server; the caller owns both shutdown and the Agent’s lifetime.
Route work through the returned run() for this admission check to apply. The wrapper does not serialize concurrent requests or stop a run already in progress. Add a concurrency limit and cancellation policy when your host accepts overlapping work. The host calls stop() after it stops admission and drains active work, then closes the borrowed Agent. Restart is a request to your supervisor, not an automatic replay of unfinished work.

Presets

These are configuration values. EdgeRuntime consumes its monitoring interval, watchdog timeout, and thermal/memory thresholds. It does not automatically apply maxTokens, contextWindow, memoryLimitMb, or disableFeatures to the Agent or operating system. Apply supported model/Agent options explicitly and use your process supervisor for process limits.

Features

Watchdog

heartbeat() updates the last-activity timestamp. If the elapsed time exceeds the preset timeout, the runtime emits watchdog-restart, increments watchdog_restarts, and invokes onWatchdogRestart when supplied. The event name reports a requested restart; no Agent replacement occurs inside the runtime. Define what a heartbeat proves. The example marks run boundaries, so a long run can legitimately exceed the timeout. Choose a suitable timeout or call heartbeat() from meaningful progress milestones. An unconditional timer can hide a stuck work loop. An idle service also needs an explicit liveness strategy if the watchdog remains enabled. Your supervisor owns cancellation/drain, Agent cleanup, process replacement, and reconciliation of any external effect whose result is unknown. See recovery.

Resource monitor

CPU thermal and memory warnings mark the runtime as degraded. When monitored CPU/memory values recover, it emits recovered. The resource monitor also reports disk and GPU warnings through runtime.getMonitor(); those are not automatic runtime restart or throttling policies.

Health endpoint

After the host starts the runtime, inspect:
The result includes state, uptime_ms, watchdog_restarts, resources, and degraded_reason. resources contains the latest snapshot; temperature and GPU measurements can be unavailable. The built-in server returns 200 for both running and degraded, and 503 for stopped while reachable. Its status is not a readiness guarantee for your application. Implement readiness from your admission policy when a load balancer must stop sending work. The built-in endpoint has no authentication or configurable bind host; protect it with your network configuration, or set disableHealthCheck: true and expose getStatus() through your existing protected server. Monitoring timers and the health server are unreferenced; they do not keep an otherwise idle Node process alive.

Config

GPU monitoring

ResourceMonitor probes nvidia-smi. If it is unavailable or fails, gpu is absent. Reading one snapshot is enough for a one-time inspection; start() and stop() add periodic events.

GPU snapshot fields

The monitor emits gpu-warning when GPU memory usage exceeds its memoryThreshold ratio, default 0.85. Its payload contains memory used/total and utilization; it is not a full resource snapshot.

Acting on GPU pressure

Use measurements in an explicit host admission rule. This factory starts a monitor and returns the gate and its cleanup function:
This policy denies GPU-dependent work when no measurement is available. Choose a different fallback for CPU workloads. The host must consult the gate before submitting work; the monitor cannot intercept runs by itself. Forward selected readings to your metrics sink explicitly—resource monitoring does not automatically create an exporter or Prometheus integration.

Verify the host behavior

Test admission and supervisor callbacks with controlled events before connecting hardware. On the target device, confirm real readings, heartbeat milestones, readiness behavior, and shutdown. Continue with cloud synchronization when your receiving service is ready.