Yamak popped up on Hacker News again this week, sandwiched between two other open-source browser agent launches. I keep a loose tally and the count for 2026 is around fifteen new projects, which feels absurd for a category that barely had a name eighteen months ago.

But scrolling through the repos, something jumped out: every single one makes a fundamental architectural choice in its first few hundred lines of code, and that choice determines what happens to your data, your logins, and whether you’ll survive the setup process. There are only three real approaches. And most coverage of these tools glosses right past the differences.

Clean room, no sessions

Managed cloud browsers are the most common pattern. Skyvern, OpenClaw’s hosted mode, OpenAI Operator all work roughly the same way: spin up a headless Chrome instance on a remote server, do the task there, send back results. Nothing to install. And nothing to work with, either.

Because the agent’s browser is a blank slate. No cookies, no saved passwords, no sessions. So when a task requires logging into Gmail or your company’s expense portal, the agent either needs your credentials shipped to their server or it hits a wall. And it will hit a wall, because modern web authentication was specifically designed to stop automated access from unfamiliar environments — CAPTCHAs, device fingerprinting, MFA prompts, all of it aimed squarely at headless browsers running in data centers.

I wrote more about this in a previous post. Cloud browsers work for tasks against public pages. But for anything behind a login, they’re fighting the architecture of the web itself.

What happens when a remote server drives your Chrome

Chrome DevTools Protocol is a debugging interface baked into every Chrome installation that lets external programs control the browser: click elements, read pages, inject scripts, take screenshots. Some browser agent projects use CDP to connect a remote server to your local Chrome. And this is the “Browser Relay” concept showing up in OpenClaw’s documentation and a handful of newer tools.

And the pitch sounds like it solves everything. The agent uses your real browser with your real sessions, so it can access anything you’re already logged into, while the intelligence runs on a powerful cloud server. No authentication headaches, no blank-slate problem.

But in practice it means a remote server can see and control everything happening in your browser, including every page you visit, every form field you fill out, and every password manager autofill popup that appears while the agent is doing its work. Your browsing data flows through their infrastructure not as a bug but as a core requirement, because CDP screenshots and DOM snapshots are how the remote agent observes the world. And the security surface is enormous. Most projects are honest enough about this in their documentation, though “somewhere in their documentation” is doing a hell of a lot of work in that sentence.

Browser Relay also introduces reliability problems that neither of the other two architectures deal with. If your Chrome restarts, or you navigate away from a tab the agent was using, or your laptop sleeps, the CDP connection breaks and the agent loses context mid-task. So you get session access at the cost of fragility.

Nobody mentioned extensions

I scrolled through 200+ comments on the Yamak thread. Not one person brought up browser extensions as an alternative architecture.

Just run inside the browser

The third approach is a Chrome extension that lives in the browser itself. No remote server puppeteering your tabs. No cloud browser starting from scratch. The extension reads the page you’re looking at, sends relevant context to whichever LLM you choose with your own API key, and acts on the response locally.

Dassi works this way. Side panel, your browser, your sessions, your key. Page content never touches intermediary infrastructure because there is no intermediary. And setup is “install from Chrome Web Store,” which starts to feel remarkable when you compare it against seven Docker commands and a YAML file.

And this architecture has real constraints. It can’t run tasks while your laptop is closed. Or parallelize across multiple browser instances. But for what most people actually need, which is doing something useful with the page they already have open, those limitations do not come up much.

Your data has to go somewhere

Every browser agent on Hacker News this month falls into one of these three buckets. Managed cloud browsers trade your session state for zero-install simplicity. Remote CDP preserves sessions but hands your browser’s steering wheel to someone else’s server. And browser-native extensions keep everything local at the cost of some flexibility nobody was using anyway.

I do not think any single architecture is universally correct. But most people evaluating AI browser agents are comparing features and benchmark scores when they should be asking something simpler: where does my data actually go, and who can see my screen?