Another AI Browser Agent Leaderboard just hit the Hacker News front page, and the top comment is someone asking whether any of these agents can actually fill out their company’s janky expense reimbursement form. Skyvern 2.0 scored 85.8% on WebVoyager. Impressive number. But I’ve been using browser agents daily for months now, and the gap between what leaderboards test and what I actually need an agent to do has been gnawing at me.

WebVoyager is a standardized test for unstandardized work

WebVoyager runs 643 tasks across 15 websites. Book a flight. Find a product. Navigate to a page. These are reasonable tasks, and scoring well on them requires genuine capability. I do not want to dismiss that.

But my actual browser work looks nothing like this, and I suspect yours does not either. Last week I needed to pull invoice data from three different vendor portals, each with its own bizarre login flow and session timeout behavior, then cross-reference those numbers against a spreadsheet someone shared with me in Google Sheets. The week before that I spent twenty minutes drafting a reply to a client email while simultaneously checking their LinkedIn profile to remember what project we last discussed together, which is the kind of multi-tab context-juggling that no benchmark captures because it’s too messy and personal to standardize.

So the leaderboard tells me Skyvern can navigate Booking.com. Great. Can it handle the procurement portal my company’s IT department built in 2019 and hasn’t touched since?

The scoreboard incentive problem

Goodhart’s Law is already doing its thing. When every browser agent company knows the WebVoyager task set, they optimize for it. Not maliciously, just naturally. You tune what you measure. And once three or four competitors cluster at 84-88%, the scores stop meaning much for differentiating between them.

I saw a Reddit thread in r/ChatGPTCoding where someone pointed out that two agents with nearly identical WebVoyager scores performed wildly differently on internal HR tools. One breezed through a PTO request form. The other got stuck in a loop clicking the same dropdown over and over. Same benchmark tier. Completely different real-world experience.

What actually matters when you pick one

I’ve been through enough browser agents at this point to know what separates the useful ones from the impressive-on-paper ones. And it comes down to three things that no benchmark measures.

The first is page context. Most benchmarks start the agent on a blank page and say “go here, do this.” But that is not how people work. You are already on a page, already mid-task, already staring at the thing you need help with. An agent that can read the page you’re on and act from there is fundamentally more useful than one that needs you to describe everything from scratch, even if the second agent scores higher on WebVoyager.

The second is model flexibility. A locked-in agent using one specific LLM might score well on the leaderboard, but if you want Claude for sensitive drafting and GPT-5.2 for research and Gemini for long documents, you need something that lets you choose. Dassi lets you bring your own key or just log in with your ChatGPT subscription, so you pick the model that fits the task instead of being stuck with whatever the top-scoring agent locked itself into.

And the third is data routing. Benchmark tests run on public websites with fake data. Your browser has real email threads, real financial documents, real client information. An agent that funnels everything through a third-party server to hit a high benchmark score might be exactly the wrong choice for your actual workflow.

The crap nobody benchmarks

Nobody is scoring agents on “can it draft a follow-up email that sounds like me after reading a thread I’m already looking at.” Nobody is testing whether the agent knows when to stop and ask before clicking “Submit” on a form that charges your corporate card. Nobody is measuring how well it handles the seventh tab you have open, the one with the half-finished Google Doc you keep switching back to.

Browser work is contextual, personal, and weird. My daily usage of dassi involves tasks that would be nearly impossible to standardize into a benchmark — summarizing a page I landed on from a Slack link, pulling specific numbers out of a dashboard that only my team has access to, drafting replies to emails where tone matters more than content.

Scores will converge. Context won’t.

A year from now, every serious browser agent will score above 90% on WebVoyager or whatever benchmark replaces it. The underlying models are getting better fast enough that raw capability on standardized tasks will stop being a differentiator, the same way raw coding benchmark scores already barely distinguish the top LLMs from each other.

What will still matter is whether the agent lives where you work. An agent that sits inside your browser and sees the page you’re already on has an architectural advantage that no amount of benchmark optimization can replicate from the outside. That advantage gets bigger, not smaller, as the models improve — because the bottleneck was never intelligence. It was context.

The leaderboard is fine. Look at it. Just don’t mistake it for a product review.