Skyvern 2.0 Scored 85% on WebVoyager. Your Logged-In Tabs Don't Care.
I saw Skyvern 2.0 trending on Hacker News this morning with a Show HN post claiming 85.8% on WebVoyager. The number is real. The benchmark is credible. And the gap between that score and what happens when you try to use a browser agent in your actual workflow is enormous.
WebVoyager is a well-known benchmark for browser agents. It spins up fresh cloud browsers, gives each one a task, and measures whether they complete it. So the tasks are real: book a flight, find a Reddit thread, search a job board. But the conditions are sterile.
You are not sterile. Your browser has 47 tabs open. Your Gmail is logged in. Your Notion doc is halfway through a draft. Your Figma file has unsaved comments. And that’s the environment where real work happens, not the environment WebVoyager tests.
Why the score doesn’t transfer
A benchmark is a controlled experiment. So controlled experiments strip variables out. And WebVoyager strips session state out, because if every agent got to start from different logged-in environments, you couldn’t compare them.
So Skyvern’s 85.8% means exactly this: on a fresh cloud browser, with no cookies, no sessions, no personal data, no authenticated state, and no memory of what you were doing five minutes ago, it completes 85.8% of benchmark tasks.
That’s impressive engineering. But it’s also a sentence that reveals why cloud-based browser agents hit a wall the moment you try to use them for real work. The moment you want the agent to open your Gmail, the demo falls apart. Because the agent is running in some datacenter. And your session cookies are in Chrome on your laptop. So there is no bridge.
A sentence worth rereading
It completes 85.8% of benchmark tasks. Benchmark tasks. Not your tasks.
The part nobody benchmarks
Almost every browser agent benchmark, including WebVoyager, tests the agent’s ability to navigate unfamiliar websites from scratch. Which is the hardest version of browser work. But it’s also the version almost nobody does day-to-day.
Most of my browser work lives in six tabs I already have open. Gmail, calendar, a CRM I’m logged into, a spreadsheet I’ve been editing for a week, the docs for the project I’m on, and whatever Claude tab I keep beside them. So the agent that helps me most is not the one that can boot a new Delta.com session and book a flight. It’s the one that can read the email I’m staring at and draft a reply in my voice.
WebVoyager doesn’t score that. And there’s no benchmark for it. Because if there were, the leaderboard would look completely different. Every cloud-first agent would score zero.
Where the gap closes
Dassi is a Chrome extension. It runs inside the tab you’re already in. So when you ask it to read your Gmail, it reads your Gmail, because you’re logged in. And when you ask it to extract data from a page behind your company’s SSO, it does that, because you’re already past the SSO wall. No OAuth handoff. No cookie cloning. No remote CDP connection to a headless browser running in someone else’s cloud and pretending to be you.
This is not a benchmark-beating setup. But it’s a session-beating setup. And for the work most people actually do, session state is the thing that matters.
If you want the long-form version of this argument, I wrote it up in Cloud Browser Agents Can’t See Your Tabs. Or go read Three Ways AI Agents Connect to Your Browser for the different architectures and what each one actually costs you in practice.
What the 85.8% is actually useful for
I’m not dunking on Skyvern. It’s open source, it’s fast, and 85.8% is a legitimate result on a genuinely hard benchmark. So if you want to automate workflows where the agent needs a clean browser anyway (scraping public data at scale, testing your own site from the outside, running flows on throwaway accounts where you can safely hand credentials over), Skyvern is a serious tool and I’d reach for it without hesitation.
But don’t read the HN headline as “browser agents solved.” Read it as “browser agents solved a specific slice of the problem, in a specific environment, under specific conditions.” The slice where the agent has to act inside your authenticated browser context is untouched by that score. And that’s the slice where Dassi lives. Free Chrome extension, log in with your ChatGPT account or bring your own API key.
So the next time a benchmark number trends, my question’s going to be the same: what did it measure, and did it measure you?