· Clane AI · 4 min read
Why your agent needs its own browser
A page fetch is not a browser. What an agent loses without sessions, a vault and a person in the loop, and what a real browser per user costs to run.
Most agents meet the web through a fetch: download the HTML, strip the tags, hand the text to the model. That works for a public article. It stops working the moment the task is something a person actually does on the web: check an order, download last month's invoice, change a setting, answer a message. Those pages sit behind a login, run on JavaScript, and expect the same browser to come back tomorrow.
We built WebPilot.si because we kept hitting that wall. This post is about why the answer is a real browser that belongs to one person, and what that costs.
A fetch forgets everything
A fetch has no memory. Every run starts as a stranger: no cookies, no local storage, no open tabs. So the agent logs in again, which means it needs the password, which means the password ends up in a prompt, a config file or a tool result. Then the site sees a new device, asks for a code, and the run stops.
A headless browser in a throwaway container is better at JavaScript but has the same amnesia. It also looks like what it is. Pages that serve a person without complaint often answer a fresh headless client with a challenge.
What the agent needs is closer to what you have on your own laptop:
- A profile that persists. Cookies, storage and open tabs survive disconnects and restarts. In WebPilot.si each user has their own Chromium process and profile; the browser stops after 20 minutes without activity, and the profile and tab list come back on the next request.
- A normal browser. Headed Chromium on its own virtual screen, with no automation flags. Sites see a browser with history and cookies, not a fresh automation session.
- The same browser for every client. Claude Code in the morning, a Python script at night: both reach the same browser with the same token, because the user comes from the token, not from the client.
Credentials the agent never sees
Once the browser remembers, the next question is how it logs in the first time, and again when a session expires. The answer we settled on is that the agent never handles a secret at all.
A person stores the login in a vault: username, password and, if the site uses an authenticator app, its setup key.
Values are encrypted with AES-256-GCM and can be written but never read back through any API. The agent calls
browser_login{site: "github.com"}; the server decrypts the values inside the gateway and types them into the page.
If the entry has a TOTP key, the server computes the code and fills that too. The agent sees a status, never a value.
This matters beyond tidiness. An agent reads untrusted text all day. If a page can talk it into printing a password, and the password is in its context, it will eventually happen. If the password was never there, it cannot.
Guardrails the server enforces
The same reasoning applies to writes. "Please don't submit anything" in a prompt is a request, not a control. So the rules live in the server:
- Read tabs block every write at the network level. Looking around can never submit a form.
- Site rules per user: a bank set to
read_onlyrefuses writes, logins, secrets and JavaScript withsite_read_only, while a shop stays writable. - Audit log: every write, login, secret use and hand-off, with the user on every line.
- Untrusted data: page text reaches the model fenced and labelled as data, and our agent instructions say never to follow it. That one is a safeguard, not a guarantee, which is why the other three exist.
A person in the loop
Some steps are not the agent's to take. A code sent by SMS. A CAPTCHA its model cannot read. A payment you want to confirm yourself. For those the agent creates a hand-off: a link to a page with its message, your browser live, and a Done button. You open it on your phone, do the step, press Done, and the agent continues where it was.
The live view is also how you keep an eye on things. Ask for a viewer link and you see the agent's browser as it works. Links are short-lived tickets, and they are view-only unless your settings allow taking over. The server's relay enforces that: on a view-only link it forwards only what is needed to draw the screen, so even a hand-made client cannot click or type.
What it costs
A real browser per user is not free, and we would rather say so. The numbers below are from the open-source server's own measurements, with the Chromium in its image:
| What is open | Memory (proportional set size) |
|---|---|
| One tab of example.com | about 250 MB |
| Wikipedia, GitHub and BBC News as well | about 610 MB |
| The user's virtual screen (Xvfb and VNC) | about 100 MB more |
The server budgets about 700 MB per browser, stops idle ones, and refuses new browsers with 503 browser_capacity
(and a Retry-After) rather than letting a host slow down for everyone. That is why our plans count
browser-hours, the time a user's browser is running, and not tool calls.
Try it
The hosted service at webpilot.si runs the same engine as the open-source server on GitHub. Connect Claude Code with one command:
claude mcp add --transport http webpilot https://api.webpilot.si/mcp \
--header "Authorization: Bearer $WEBPILOT_TOKEN"Then ask it: "Open example.com in the browser, tell me the heading, and give me a viewer link."