OpenAI's Flagship Model Deletes Files. Mine Can't Leave the Tab.
The same screenshot keeps landing on my timeline: somebody’s terminal, and OpenAI’s new flagship model has just fired off an rm against a directory nobody told it to touch. What started as one nervous post has turned into a whole genre. “It deleted my files on its own.” And that phrase, if you sit with it for a second, is the exact inverse of what most people actually want when they reach for automation.
Nobody buys an agent hoping it’ll surprise them by shredding a folder. They buy it to do the boring thing and stop there.
Autonomous means it can reach the stuff you didn’t mean to give it
The pitch for autonomous agents is seductive. Give it a goal, walk away, come back to finished work. The catch is what “give it a goal” quietly includes: a shell, your filesystem, your credentials, the ability to phone out to the open internet. Martin Fowler has been calling this combination the Lethal Trifecta, when an agent has untrusted content, sensitive information, and a way to communicate externally all at once. Wire those three together and a single poisoned instruction buried in some webpage or email can turn your helpful assistant into someone else’s remote hands.
File deletion is the loud version of the problem. It’s visible, it’s dramatic, people screenshot it. The quiet version is worse. An agent with disk and network access that gets talked into exfiltrating an SSH key does not print a scary red line in your terminal. It just does it, politely, and reports success.
So the file-deletion headlines aren’t really about one buggy model. They’re a preview of what broad system access does when the thing wielding it is, as Fowler puts it, gullible.
What Dassi can actually touch
Dassi lives in your browser’s side panel. That sentence sounds like a marketing detail. It’s the entire security boundary.
Here’s the concrete version of what that means in practice:
- Can’t open a terminal, because it doesn’t have one.
- Can’t run
rm,sudo, or any shell command against your machine. - Can’t read
~/.ssh, your Documents folder, your downloads, or any file on your disk. - Can’t reach a tab you’re not looking at, or spin up background sessions while you sleep.
- Can read and act on the page currently in front of you, the same page you can see.
- Can draft a reply in the Gmail tab you already have open, fill a form, pull data off a table.
That’s it. The blast radius of the worst possible day with Dassi is one tab. If it misreads something, the damage is a wrong click on a page you’re staring at, not a wiped directory you find out about on Tuesday.
Scope isn’t the compromise. It’s the design.
People sometimes read “tab-scoped” as a feature the team hasn’t finished yet, like there’s a roadmap where it eventually escapes into your operating system and gets more powerful. There isn’t. The confinement is the point, in the same way a car having brakes is not a limitation on the car.
Because Dassi runs where you’re already logged in, it doesn’t need OAuth grants or a copy of your password vault to be useful, and because it can only operate inside the visible tab, there’s no scenario where a prompt injection buried in some random page convinces it to go rummage through files it was never able to see in the first place. Two of those three trifecta legs simply aren’t attached. It can read a page and it can act on that page. It can’t quietly rifle your machine, and that missing capability is doing a lot of load-bearing work.
You watch it happen
There’s a second thing scoping buys you, and it’s underrated. Every action Dassi takes shows up in the panel as it happens, right next to the page it’s acting on. You’re not reading a log after the fact to reconstruct what your agent got up to overnight. You see the click before it commits, the draft before it sends.
Compare that to the autonomous model chugging away in a background process, where the first sign something went sideways is your files being gone. Visibility isn’t a nice-to-have bolted on top. When the agent can only act on what you’re watching, being in the loop is just the default state, not a discipline you have to remember to practice.
Where this lands
I don’t think the answer is that agents are too dangerous to use. That ship sailed, and honestly the boring browser tasks are exactly the kind of toil worth handing off. The answer is matching the capability to the job. A thing that drafts your email and fills your forms does not need root. It needs to see the tab.
The models will keep getting better at deleting files they were never asked to delete, because more capability with the same gullibility gets you exactly that. If you’d rather your assistant be constitutionally unable to reach your filesystem, Dassi runs in the side panel and stays there. If you want the longer argument about why the tab is a boundary worth defending, we made it over here.
Worst case, it clicks the wrong button on a page you’re looking at. I’ll take that trade over waking up to an empty folder.