Gemini turned up this week as the latest frontier model named in an intrusion campaign run against other companies, and within about six hours my timeline had converged on one take: the AI safety conversations have gotten unbelievable. Maybe they have. But underneath the dunking there’s a question people are asking in earnest, and it’s much smaller than the headline: if I install an AI agent, what does it actually get to see on my machine?

I can answer that for one shape of agent. The Chrome extension kind, the one that lives in a side panel next to the page. We build one, so read this as interested testimony. The scope just happens to be unusually easy to describe, which is most of why it’s worth describing.

The tab you’re looking at, and not much behind it

A side-panel agent reads the rendered page of the active tab. Rendered, not the raw HTML your server sent — the DOM after JavaScript has run, after React hydrated, after the infinite scroll loaded the next forty rows. It also reads the accessibility tree, which is the same structure a screen reader consumes: roles, labels, states, which button is actually a button. Text, links, form fields, table cells. A screenshot when the page is visual enough that text won’t do.

That’s the input. We wrote up the mechanics of DOM versus screenshot versus accessibility tree in a separate post if you want the tradeoffs between them.

For actions: it clicks, types, scrolls, and navigates inside your existing session. Whatever Gmail shows you, it sees, because it’s looking through the same authenticated window you are.

No token, no grant, no standing credential

Here’s what I think people miss when they compare agent architectures. A cloud agent that wants to do something with your Gmail has to be given Gmail — an OAuth grant, usually scoped somewhere between “read all mail” and “read, send, delete, and manage settings,” held on a server you don’t operate, refreshable indefinitely until you remember to go revoke it. Same for Slack. Same for your CRM. Every app is a separate negotiation, and the end state is a service somewhere holding a fan of long-lived keys to your working life, each one valid whether you’re at your desk or asleep.

A browser extension has none of that. It holds no token for Gmail. It has no credential it could hand to anyone, because there’s nothing to hand over. What it has is the ability to act inside a session cookie that Chrome already holds and that your browser already renews on your behalf. Close the tab and the agent has nothing. Log out of Gmail and the agent is logged out of Gmail. Uninstall the extension and there is no dangling grant sitting in some provider’s database that you’ll forget about for two years.

So the blast radius is your open browser, during the time your browser is open, on the pages you’ve navigated to. And the revocation story is a right-click and “Remove from Chrome,” which is a damn sight simpler than auditing an OAuth console.

I’d rather have that property than a more capable agent, and I don’t think that’s a close call. dassi works this way for exactly this reason.

Where it goes blind

The honest list of things it can’t reach:

  • Cross-origin iframes. An embedded Stripe checkout, a Zendesk widget, a third-party payment frame. The browser’s same-origin policy blocks script access, so the agent can see the frame exists and can’t read inside it. Screenshots help a little. Not much.
  • Tabs you haven’t opened. There’s no background sweep of your browsing history. If it isn’t open and active, it isn’t in context.
  • Other applications. Your desktop, your files, your Slack app, your terminal. Outside the browser, outside the agent.
  • Canvas, WebGL, video. Rendered pixels with no text layer. A dashboard drawn entirely in canvas is a picture of numbers, not numbers.
  • Closed shadow DOM. Rare, but it happens, and it’s opaque by design.

Some of these are fixable with better tooling and some are the browser’s security model doing precisely its job. I’m fine with the second category staying broken.

What leaves your browser, then

Page content goes to the model. That’s the part worth being clear-eyed about, because “runs locally” and “sends nothing anywhere” are not the same claim, and anyone selling you the second one is selling you something.

Working backwards: the model has to read the page to reason about it, so the relevant text from the tab travels to whichever provider you picked. With bring-your-own-key, that trip is your browser to Anthropic or OpenAI or Google, under your own account, under your provider’s retention terms. No vendor middleman in between logging prompts. If you log in with a ChatGPT subscription instead, it’s your browser to OpenAI, same idea. We went deeper on the routing question here.

The rest stays put. Settings, history, page context between turns: local storage.

Boring on purpose

None of this makes an extension agent safer than a cloud agent in some absolute sense. Prompt injection from a hostile page is real, and a browser agent reading attacker-controlled text inside your logged-in session is a genuine attack surface that nobody has fully solved.

But the scope is legible. You can hold the whole thing in your head: this tab, this session, right now. Try saying that out loud about a system holding refresh tokens for nine of your SaaS accounts.