Desktop AI Went Local. Browser Agents Are Still Renting Servers.
Saw a Show HN this morning called Desktop Agent Center. Local AI automation triggered by hotkeys, with the whole pitch being that nothing leaves the machine. No cloud round trip, no telemetry, just an LLM running against your actual desktop apps. People in the comments were excited about the privacy story, but I noticed something else. The desktop side of agentic AI is converging on local execution while the browser side is doing the opposite.
Two timelines, opposite directions
Pull up the agent landscape from a year ago and almost everything ran in the cloud. Desktop agents had to. There weren’t good local models, the OS-level tooling didn’t exist, and shipping anything that automated your computer meant either a remote VM or a sketchy permission grant. And browser agents copied the same architecture because it was the only one available.
Now the desktop folks have figured out hotkey-driven local control. Apple’s accessibility APIs, Windows UI Automation, and a handful of Linux DBus calls turned out to be enough scaffolding for a real agent. Local models got smaller and faster. The trick Desktop Agent Center pulls off is mostly plumbing, but it’s plumbing that makes the cloud version look antique.
Meanwhile, look at what shipped in browser-agent land this quarter. Comet runs in their cloud. Operator runs in OpenAI’s cloud. Half a dozen YC companies announced “autoscaling browser fleets” and acted like that was a feature instead of a tax.
The cloud VM is clicking a Chrome that isn’t yours
Most cloud browser agents work the same way. You ask them to do something. They spin up a sandboxed Chromium on a server somewhere. That Chromium has none of your cookies, none of your sessions, none of the muscle memory your real browser has built up over years of being you, and the first time it hits a 2FA prompt the agent grinds to a halt asking for a code you have to type by hand.
The agent isn’t dumb. But the architecture is dumb. Your real browser, the one you are reading this post in, is already authenticated to the services you want to automate. It already has the right zoom on your weird high-DPI monitor. It knows which Google account is your work one and which is the throwaway you use to leave Yelp reviews. And none of that survives a fresh VM in us-east-1, which is a hell of a starting point for an agent that’s supposed to save you time.
This is the same problem we keep coming back to. Cloud browser agents can’t see your tabs and can’t see your logins. The desktop agent crowd just demonstrated, by accident, that the local approach works fine when you actually try it.
So why hasn’t the browser caught up
Partly money. Cloud agents bill per task. Local agents don’t bill at all once installed. Investors prefer the first model.
Partly inertia. The first wave of browser-agent companies built on Playwright fleets because Playwright fleets were what AI infra vendors were already selling. And the second wave inherited the architecture. Nobody stopped to ask whether it made sense.
And partly a real engineering problem. Running an agent inside the user’s actual Chrome means dealing with extension APIs, MV3 service workers, sandboxed iframes, and Chrome’s continual attempts to lock down what extensions can do. It is harder than spinning up a VM. But “harder” is not “wrong.”
This is the bet dassi made from day one. The agent lives in the side panel of your existing Chrome. Your sessions, your tabs, your cookies, your scroll position — it sees what you see. So nothing ships to a cloud Chrome because there is no cloud Chrome. The model call goes out using your own key, the model call comes back, and the click happens locally.
Local everything
Local desktop. Local browser. Same idea. Different year.
What the Show HN actually proved
The Desktop Agent Center launch isn’t going to dethrone anything. Most people will not download a hotkey-driven local agent for their Mac and rebuild their workflow around it. But the demo did one useful thing. It made the cloud-VM browser-agent architecture look strangely dated by comparison.
If a hobby project can wire up local control of arbitrary desktop apps, the question of why a Series B startup needs cloud Chromes to click a button gets sharper, and the answer is mostly about who pays the AWS bill rather than what’s technically possible. The browser is the easier surface, not the harder one. And dassi has been running local-first inside your real Chrome the whole time. It works.
The desktop side just got there a little faster, which is funny considering the browser is where most people actually do their work all day.