A post titled “Open-weight AI models are catching up to the frontier. The safety gap remains.” sat near the top of Hacker News this week, and the thread went exactly where those threads go: benchmark disputes, license lawyering, whether Qwen counts as open if the data isn’t. Almost nobody fought the capability claim. That argument is finished.

The second sentence is the one worth sitting with.

The gap moved, it didn’t close

Capability and alignment get built by different teams on different budgets, and only one of the two shows up on a leaderboard. A lab six weeks out from a release that needs a competitive agentic score will spend its remaining compute on that score, because refusal behavior and instruction-hierarchy hardening and jailbreak resistance are slow, unglamorous, and completely invisible to everyone who’s going to write the launch post.

So you get a model that plans a nine-step web task correctly and also obeys a paragraph of text buried in the fourth reply of a support-forum thread.

In a chat window that’s a curiosity. Give the same model a browser and it’s a liability.

Same day, a pile of agents running where nobody looks

Scroll down that same front page and there they were: Yamak, autoscaling browser-agent fleets, a couple of fresh leaderboards. Every one of them spins up a headless Chrome in a datacenter, runs your task, and hands back a summary.

The summary is the product. The run is not.

That’s the part I don’t think gets priced in. When each step executes in 300ms on a machine you will never open a terminal into, “supervising the agent” collapses into reading a transcript the agent wrote about itself. I’ve made this complaint about Yamak specifically, but it isn’t really about Yamak. It’s the default shape.

Which is a strange thing to hand a model with a documented compliance problem

You wouldn’t give a first-week contractor production credentials and a “tell me how it went” review cycle. Somehow that is the standard deployment for agents right now.

Put it in a tab

The fix is geometry, not virtue. Keep the execution somewhere your eyes already are.

Dassi is a Chrome side-panel agent, so the model drives the tab sitting next to it. You watch the click land. You watch the form fill. When it opens something you didn’t expect, you see the page load before the next action goes out, and closing the panel or hitting the tab is a physical interrupt rather than an API call to a queue.

Because it’s BYOK, the model underneath is yours to pick:

  • DeepSeek — cheap, strong at multi-step planning, fine for bulk extraction
  • Kimi — long context, good on dense pages (more on running it here)
  • Qwen — solid at structured output, worth trying on form-heavy work
  • Or GPT/Claude/Gemini if the task is one where you want the alignment tax paid for you

Paste a key, pick the model, run the same task twice with two different ones. Takes a minute. The extension is here.

Watching isn’t a guarantee, it’s a speed limit

I’m not going to pretend a human in the loop is a security control. Attention decays. After the fortieth approval you’re clicking through on muscle memory, and anyone who says otherwise has never done data entry for a living. Prompt injection can still land, and a model that’s mid-task on a page you trust can be handed instructions by content you never read.

What changes is the blast radius and the clock. A cloud agent with your credentials in its environment can execute forty actions across six sites before the run report exists. The same model in your side panel gets one tab, one action, and however many seconds pass before you look up — and those seconds are the entire difference between “it did something weird and I stopped it” and “I found out on Thursday.” The exposure is bounded by what a single visible browser tab can reach, which is a much smaller thing than a fleet with a credential vault.

And the data path stays narrow. Your key, your provider, no middleman logging the run, which matters more now that open-weight inference is scattered across a dozen hosts of wildly varying seriousness. I’ve written before about why BYOK is the only routing story that holds up.

The models got good enough that the alignment gap stopped being an academic footnote and started being a deployment question. My answer is to keep the damn thing in front of me until the safety work catches up, which, judging by how these release cycles have gone, gives me a while.