OpenAI Built a Lockdown Mode for Prompt Injection. That's a Confession.
OpenAI shipped something called Lockdown Mode last week, and the name is doing a lot of work. Lockdown. As in: things got bad enough that the answer is a panic room. The feature is meant to protect sensitive data from prompt injection, the attack where a hidden instruction buried in content your agent reads quietly hijacks whatever it does next.
I think the feature is fine. What’s interesting is what it admits.
A defense mode is a confession
When you ship a special hardened mode, you’re admitting the normal mode has a problem you couldn’t design away. Lockdown Mode exists because ChatGPT’s agent fetches web pages, documents, and tool outputs through OpenAI’s own servers, and any of that content can smuggle in instructions the model will follow like they came from you. The untrusted text and your sensitive data land in the same context window, on infrastructure you don’t control.
So the fix is to clamp down. Restrict what the agent can touch, narrow what it can send out, wall off the sensitive material. Sensible enough. But it’s a patch bolted onto an architecture, not a property baked into one.
The trifecta nobody can dodge
Simon Willison has been hammering on this for more than a year with what he calls the “lethal trifecta.” The idea is simple. An agent that holds access to your private data, plus exposure to untrusted content, plus the ability to communicate with the outside world, is a prompt injection bomb sitting there waiting for the right match to come along. Pull out any one of those three legs and the bomb mostly defuses, which is exactly the move Lockdown Mode makes when it restricts what the agent can reach and where it can send things.
Cloud agents, though, are built to hold all three at once, because all three together is the product people pay for. They read your private stuff. They browse the open web on your behalf, which means they’re constantly swallowing text written by strangers who don’t have your interests at heart. And they reach out to APIs, inboxes, and external services to actually get things done. Lockdown Mode is OpenAI reaching into its own architecture and sawing off one leg when the danger spikes, then bolting it back on the moment you want full capability again.
Martin Fowler described the same fear about agentic email, where someone he talked to boxed the agent in with read-only access, no internet, and drafts written to a plain text file for human review. The capability drops through the floor. That’s the toll for standing outside the trifecta.
Where the content actually goes
The lockdown framing skips a step. The real danger isn’t only that the agent reads a poisoned page. It’s that the page gets fetched and processed somewhere far away from you, in a session you can’t watch, using credentials and context you handed over once and then stopped thinking about entirely, the way you forget a spare key you left with a neighbor.
A browser agent flips that geometry. dassi runs as a Chrome extension that lives in the side panel of the tab you already have open, and it acts inside your existing session, the one where you’re already logged into Gmail or your CRM or your bank. Nothing routes through a separate cloud browser that crawls the web on its own and ships the results back to a server farm.
That doesn’t make prompt injection vanish. A malicious page is still a malicious page. But the blast radius shrinks by a hell of a lot, because the agent works in one tab, on one task, with you sitting right there, instead of holding a master key to a cloud session that quietly touches everything you own.
You’re sitting right there
When the work happens in your side panel, you watch it happen. The agent clicks, you see the click. Something reads wrong, you stop it before the next step.
What I’d rather see
I don’t buy hardened modes as the future. They feel like seatbelts welded onto a car that was designed to drive off cliffs. The more durable answer is to keep the agent somewhere the trifecta never fully assembles: in your browser, in your session, under your own eye, sending things only where you can watch them go. A person glancing at the screen catches the weird stuff no policy filter on earth anticipated, and that supervision comes for free when the agent works in your real tab instead of a datacenter you’ll never see the inside of. We made that case earlier about why watching your tab beats inventing a protocol for invisible agents.
OpenAI knows the attack surface is structural. They shipped a whole mode to prove it. You can try dassi from the Chrome Web Store and keep the work in the one place you’re already looking. Whether the next wave of agents keeps building the panic room or just stops moving into the neighborhood that needs one, I can’t tell yet.