dassi Scores 85% on Odysseys With the Official Scorer, on a Flash Model
When Carnegie Mellon published Odysseys in April, the number I kept rereading was 44.5%. That was Claude Opus 4.6, running as a computer-use agent with a 100-step budget, fully finishing fewer than half of 200 web tasks that real people had done in their own browsers. (Lawrence Jang, Jing Yu Koh, Daniel Fried and Ruslan Salakhutdinov rebuilt the tasks from actual Chrome histories, which is why I find them more interesting than most benchmarks.)
We ran dassi through all 200 on Gemini 3.8 Flash, the model you get by default, and graded the runs with the authors’ own scorer and judge. It fully completed 170 of them, 85.0%.

We first reported 91.5%. That number came from our own judge, and it was too generous, so we took that post down and regraded everything with the official code. The rest of this post is the corrected picture.
What “fully completed” means
Every Odysseys task comes with between 3 and 12 rubric items, 1,225 across the set, and they read like a demanding client wrote them: confirm that two Kohl’s gifts can be picked up in store today, report the shop name and the customization options on an Etsy listing, and so on. A task only counts when every item passes. Get eight of nine checkpoints right on a long errand and the headline number gives you nothing.
So 85.0% means 170 tasks where nothing was missed. Across the individual rubric items, dassi passed 1,142 of 1,225, about 93%.
Why our judge said 91.5 and the official one says 85
The authors’ scorer is stricter in three specific ways, and I went through every rubric item it failed on a task our judge had called perfect. There were 17 of them across 14 tasks.
- Blocked sites count as failures. Seven of the 17 were sites that stopped the agent: a CAPTCHA on Skyscanner, a PerimeterX check on The Kitchn, reCAPTCHA on U-Pack’s quote form, an Access Denied page on Theory. The agent tried to get past each one, said plainly that it couldn’t, and found the information somewhere else. Our judge gave credit for that. The official rules say a blocked agent fails the item, full stop.
- It grades what the trajectory shows, not what the agent concludes. Another seven were cases where the agent decided the thing didn’t exist, like “Hertz has no Honda in Barcelona”, or came back with fewer fields than asked for. The official judge wants that grounded in the pages the agent actually visited.
- End states have to be visible. The last three were requirements like “keep the four strongest listings open in tabs”. Our judge let the agent’s word stand; the official one wants to see it.
There’s a fourth thing I’d rather flag than argue. Our runs happen in headless Chrome on a CI server, behind a single residential proxy. That setup draws bot checks a person’s own logged-in Chrome often wouldn’t, and a couple of the blocks (a site that wouldn’t connect through the proxy at all, for instance) look like the environment more than the site. I can’t prove how much of the gap that explains, and every agent on the leaderboard is graded under the same rule, so the official number is the one we’ll use.
How it compares
The leaderboard entries don’t all run under the same step budget, so here is dassi at each one.
- No step limit: 85.0%. BrowserCode on GPT-5.6 Luna, which averaged 124 steps a task, sits at 86.0%.
- 200 steps: 84.5%. Aside on GPT-5.5 reports 75.5% at the same budget.
- 100 steps: 75.0%. Skyvern on Claude Opus 5 reports 90.5% there, and it’s the clear leader.
The budget comparison flatters nobody, including us. A dassi step is one model call, and one call can run a small program that clicks through a whole page, so 100 dassi steps and 100 screenshot-and-click steps aren’t quite the same unit. Skyvern averaged 65 steps and still got there first, though, and I don’t want to explain that away.
The hardest tasks
The authors split the set by difficulty, and 109 of the 200 tasks land in “hard”: the longest errands, spread across the most websites.

dassi fully completed 94 of the 109 hard tasks, 86.2%, slightly more than its overall rate. The computer-use models in the paper go the other way: Claude Opus 4.6 drops from 44.5% overall to 11.0% on the hard split, and GPT-5.4 to 3.7%. When a task means comparing products across four retailers and leaving the right pages open, an agent that guesses where to click on a screenshot dozens of times in a row piles up small mistakes until one of them sinks the whole thing. dassi reads each page as structure (elements, labels, state) instead of pixels, which I think is most of why it holds up on the long ones.
Checking our work
This is dassi as it shipped at the time, release 0.74.1, on its default model, one attempt per task, no website logins and no per-site tuning. The median task cost about $0.66 in model calls and took around six minutes. Every task’s answer, full session, our old verdict and the official scorer’s result are in the evals repo, along with the script that feeds our runs to the official scorer, so anyone can rerun the grading.
Getting marked down by the official scorer is a bit humbling. It’s also the number worth quoting, and the 30 tasks we missed are a pretty good to-do list, starting with the ones where a CAPTCHA got the last word.