Three Browser Agents Hit HN Today. None of Them Get Past Okta.
Yamak launched on Hacker News this morning. So did an open-source autoscaling browser agent, and an AI Browser Agent Leaderboard ranking a dozen of them by accuracy. I read all three writeups looking for one specific thing in the task lists: anything behind a corporate login. Booking flights, scraping GitHub issues, filling public forms, comparing prices on Amazon. Every benchmark task on every list runs on the open web.
That isn’t sloppy benchmark design. It’s the actual boundary of what these things can do.
Step one is already the end
Point a cloud agent at jira.yourcompany.com and watch the network trace. First request returns a 302 to yourcompany.okta.com, and the agent lands on a username field it has no username for. Most demos solve this by handing the agent credentials, which is the moment somebody on your security team starts drafting an incident report, because SSO credentials sitting in an agent’s config file are SSO credentials sitting in a config file regardless of how they got encrypted on the way in.
But say you do it anyway. Password accepted. Okta now wants a second factor.
WebAuthn does not travel
If your company standardized on security keys or platform authenticators, the private key is bound to hardware on your desk. Non-exportable by design. That is the entire point of the spec, and it is a good property that I do not want anyone to weaken. A Chrome process running in a container in us-east-1 cannot produce that assertion. Model quality is irrelevant here. Opus 5 does not have your fingerprint.
Push MFA fails differently. It works, technically, in that someone’s phone buzzes and someone taps approve. So your automation now depends on a human approving a login they did not initiate, which is precisely the behavior every security awareness training tells people to refuse.
The IP is wrong before anyone types a password
Conditional access checks where the request came from. Your agent’s browser sits on an AWS or Fly datacenter range. Okta flags it, geo-velocity rules fire, and the session gets challenged or killed.
Nobody benchmarks this because nobody can
Layer on device trust and it gets bleaker. Okta Verify, Chrome Enterprise device attestation, and MDM-issued client certificates all ask the same question: is this an enrolled machine belonging to an employee? A freshly spawned cloud VM answers no, correctly, every single time. And there is no clever prompt that turns a no into a yes, because the check is cryptographic and happens below the layer the agent operates on.
Then the part people forget. Even in the fantasy where all of that passes, the session cookie lives in an ephemeral container that gets torn down when the task ends, so the next task starts the whole gauntlet over, which means the per-task cost of an “authenticated” cloud agent includes a full SSO round trip with a human in it. I wrote about the wider version of this in cloud browser agents can’t see your tabs, but SSO is where the abstract complaint turns into a hard stop with an error page.
So the leaderboards measure what is measurable. Public sites, no auth, reproducible across runs. The results are real, and I don’t think the people building these are being dishonest. They’re just benchmarking the ten percent of a knowledge worker’s browser day that doesn’t sit behind a login. Your Jira, your Zendesk queue, your internal admin panel, your Workday, your Snowflake console, the finance dashboard that only exists at a .internal hostname. None of it is on any leaderboard because none of it can be.
Your Chrome already walked through that door hours ago
You logged into Okta at 9am. You tapped the push, the device certificate checked out, your IP was your office or your home, and Chrome wrote a session cookie that is good for the rest of the day.
An extension running in that profile inherits all of it. Not by stealing anything, not by holding your password, and not by pretending to be a device it isn’t. It just runs in the browser where the work already happens. That’s how Dassi works, and the reason it can read an internal ticket queue is boringly unglamorous: the tab was already open and already authenticated.
Same models, by the way. Bring your own key, or sign in with the ChatGPT subscription you’re already paying for. The difference isn’t intelligence. It’s which side of the login wall the intelligence is standing on, which is roughly the point I made about logged-in state a while back and keep having to make again.
I’d love for someone to add an authenticated-app category to one of these leaderboards. It would be a very short results table.