Matt Ronge posted about Astropad Workbench on Tuesday, pitching it as “remote desktop reimagined for AI agents.” I’ve been turning this over for a couple days and keep coming back to the same reaction: streaming pixels to a machine that can read structured data is solving the wrong layer entirely.

A technology built for two humans on a phone call

Remote desktop exists because a support tech needed to see exactly what a confused user was seeing. Same rendering, same layout, someone saying “click the blue button, no, the other one.” Both ends of that connection had human eyes, and the pixel-faithful mirroring was the whole product.

But an AI agent is not a human squinting at a screen. So when you hand it a remote desktop stream, it receives a compressed screenshot, runs vision models to figure out which blobs are probably buttons, estimates pixel coordinates, and fires synthetic mouse events back across the wire. Every single one of those steps discards information that already existed in structured form before the rendering pipeline crushed it into an image. The DOM, the labeled machine-readable tree underneath every webpage, was right there the whole time, and any browser extension can query it without vision models or coordinate guessing or any of that overhead.

Streaming pixels to an agent is like printing a spreadsheet so you can OCR it back into cells. And nobody at Astropad would do that with a spreadsheet they already had open.

Work lives in browser tabs now

Cloud-based AI agents keep failing because they start from a blank slate. So Astropad’s pitch is letting the agent see your desktop instead, and I get why that resonates with people.

But most knowledge work in 2026 happens inside browser tabs. Email, project management, CRMs, docs, analytics dashboards. All behind authentication, all carrying session state that accumulated over weeks of staying logged in. And a remote desktop stream gives you pixels of those tabs. Nothing more.

A browser-native agent running as an extension gets the full page structure plus every cookie and active session the browser already holds, which means nobody has to re-authenticate or stand up streaming infrastructure just so an agent can read a page it is already sitting inside of. Dassi runs inside Chrome’s side panel, reads the DOM of whatever tab you’re on, and inherits your login sessions directly.

Fifteen demos, not one real workflow

I scrolled through remote-desktop-for-agents demos on Twitter last month and watched maybe fifteen clips. Every single one showed a single action: agent clicks a button, crowd goes wild. And not one ran longer than two steps. Nobody films the five-minute workflow where the agent misidentifies a dropdown on try three and the whole sequence falls apart.

Because actual work does not look like a demo reel. You’re clicking through three dropdown menus, reading values from a table, copying a field to another tab, filling out a form at the end, and with remote desktop each of those actions becomes its own screenshot-analyze-act cycle with the latency compounding because every step depends on the last one completing correctly before the agent can even register what changed on screen. I spent an afternoon last week watching one of these pixel-based agents try to navigate a Salesforce dashboard, and it took four attempts to correctly identify a filter dropdown that sat right next to another dropdown with nearly identical styling. Four attempts on a single element in a twenty-step workflow.

But browser agents do not have that problem because they read element IDs, class names, aria labels, and click by CSS selector instead of guessing at pixel coordinates. Over a multi-step workflow the pixel approach gets damn near unusable, and I have yet to see a single demo that goes past two steps.

Thick clients, sure

Workbench makes sense for Photoshop, legacy apps, things that never left native. But most of what Astropad is pitching at lives in a browser anyway.