Skip to main content
@agentium/browser operates web pages through Playwright. Use BrowserAgent when a model should choose browser actions, or BrowserProvider when your application already knows the navigation and interactions. Prefer an application API/tool when the task does not require a UI.

Run your first browser task

This example reads the heading on a public example page. It requires the supported Node runtime, Chromium, an OpenAI key, and network access to both the model API and the page. It does not run offline. In a new project:
Save as read-page.ts:
read-page.ts
Expect the response to identify the page’s “Example Domain” heading. Inspect finalUrl and the recorded actions as well as the answer. Exact model wording and action count vary. BrowserAgent.run() owns and closes its per-run browser connection when it finalizes; there is no BrowserAgent.close() method in v4. Change the task to read a different fact on the same page. Then try your own permitted page, updating the starting URL and domain policy together. For a visual task, select a vision-capable model and enable vision.

Choose a planning and observation path

DOM indexes identify controls in the current observation. Refresh the observation after navigation or a meaningful DOM change; an earlier index is not a permanent element identifier. Coordinates are a fallback in CSS pixels and depend on the current viewport.

Jev planner

Host factory: supply a text model and configure TypeSafe credentials according to Jev setup. Do not pass jev() as a screenshot model. The returned BrowserAgent owns the browser connection used by each run.
Evaluate completion and selected actions for the actual page. A bounded choice set does not establish that every selected action is appropriate for your application’s domain rules.

Use direct browser operations

Host function: pass a URL the host has decided to visit. This example launches Chromium, navigates, reads the DOM, and closes the browser without a model call. It still performs page network I/O and requires the browser binary.
For tabs, use newTab(), listTabs(), switchTab(), and closeTab() on the provider. Keep the returned tab IDs rather than assuming a particular ID. Refresh each tab’s DOM before using its element indexes.

Compose a browser as an Agent tool

Host function: the parent Agent decides when to invoke a browser task. Both model calls require the supplied provider’s capabilities. This function owns the parent Agent; BrowserAgent owns the connection created by its browser run.
asTool() carries the parent identity, policy, and cancellation context. Custom ToolDef execution uses core validation, execution policy, and approval. Native browser click/type/navigation/evaluate actions are not covered by that custom-tool policy. Plan mode rejects browser execution before launch.

Configure the application boundary

This is a configuration guide, not a replacement for the exported BrowserAgentConfig, BrowserRunOpts, BrowserRunOutput, and BrowserAction types in the installed package. See the package map for matching versions.

Account state and credentials

storageState loads a Playwright state file; saveStorageState on a run writes one for a later connection. Saving state does not guarantee a login succeeded or that a future session remains authenticated. Verify the application’s signed-in state and treat the file as account data. CredentialVault substitutes selected values during actions so the prompt can use placeholders. This does not mean every page string, screenshot, recording, or tool result is redacted. Filled form values are excluded from DOM labels, while arbitrary page text and screenshot pixels can still contain sensitive data. Decide what may leave your browser environment before enabling observation or recording. Domain restrictions constrain the navigation paths checked by the runtime; they are not a complete network isolation policy for every browser subresource or page effect. allowEvaluate enables JavaScript execution in page context and is disabled by default. Local browser actions and host tool execution have different authorization boundaries.

Recorded evidence and cost

Read steps, finalUrl, finalScreenshot, optional extractedContent, and optional videoPath from the result. A step screenshot and a final screenshot serve different purposes; inspect the output rather than assuming one image proves the entire task. Supply model-keyed prices when using CostTracker; follow cost patterns. A budget estimate is not a guaranteed cap on provider billing. Step limits and a bounded workload are independent controls.

Connect a local Socket.IO interface

The v4 browser gateway is a browser-specific adapter. It does not share all the authenticated text gateway’s session ownership and cancellation contracts. Host integration: the host supplies a Socket.IO server and browser instance. Use this for a trusted local interface; bind its underlying HTTP server to loopback and close the server when finished.
In v4, the gateway does not pass its abort signal into BrowserAgent.run(). Stop/disconnect therefore does not establish that browser execution or an external effect stopped. A host needing execution cancellation should call run() with its own signal and preserve the identity/ownership boundary. The optional authMiddleware can control admission; it does not automatically isolate shared browser instances or their event streams between users.

Diagnose a failed browser task

The documentation examples are type-checked; this does not establish live model, target-site, or login success. Test against an owned development page before connecting an application account.

More patterns

Continue with browser examples, core tool patterns, sandbox execution, and shipping profiles. Browser automation, API tools, and code execution solve different integration tasks.