AI Agents Can Think. They Still Can't Click.
I spent an evening last week setting up OpenClaw, the open-source AI agent that’s been gaining traction on r/selfhosted. It handles email, manages calendars, answers questions about your documents, all without shipping your data to some third-party cloud. And getting it running was smooth until I hit the Gmail step.
OpenClaw needs a client.json file. OAuth credentials from Google Cloud Console.
The 23-click credential maze
If you’ve done this before, you know the drill. Navigate to console.cloud.google.com, find or create a project, enable the Gmail API, configure an OAuth consent screen with scopes that sound like they were named by a standards committee, create the credential, download the JSON file. About 23 clicks across 7 distinct screens, and that’s assuming you don’t get turned around because “APIs & Services” contains subsections that look nearly identical but do completely different things.
I do not use GCP often enough to remember where anything lives. So instead of spending 20 minutes relearning the layout, I opened Dassi in the Chrome side panel and told it what I needed.
Done before my coffee got cold
Dassi found the project, navigated to APIs & Services, enabled the Gmail API, worked through the consent screen configuration, created an OAuth 2.0 client ID, and downloaded the client.json.
The part that actually matters
But speed was not the win. The real friction in Google Cloud Console is not that individual clicks are hard. It’s that there are six menu items that all look like they might be the right one, and “OAuth consent screen” is a separate workflow from “Credentials” even though they seem like they should be the same thing. And the sidebar rearranges itself based on your usage history, so nothing is where you left it last time. The browser agent could parse the DOM and match page elements against my request without any of that confusion.
And this is where the story turns a little absurd. I was doing all of this so that OpenClaw could access Gmail. OpenClaw is smart. It reasons about email threads, plans multi-step workflows, drafts context-aware replies. But it runs in a terminal. No browser. When the setup process requires navigating a web-based admin console with dynamic menus and stateful UI, OpenClaw literally cannot participate in its own configuration.
ChatGPT would have given me a step-by-step guide. So would Claude. But I’d still be alt-tabbing between those instructions and a console that Google redesigned sometime in February, trying to match descriptions written against a previous version of the UI to whatever was actually on my screen.
Every damn admin panel
And I keep running into this. Stripe’s webhook configuration page. Cloudflare DNS buried three levels into a dashboard that seems to add new sections quarterly. AWS IAM, which feels like it was built to punish anyone who doesn’t live in it full-time.
All web pages. And every one of them runs in Chrome.
So last week when I needed to rotate Stripe API keys, same approach. Opened Dassi, described the task, let it navigate. Because a browser agent with full DOM access reads buttons, dropdowns, form labels, error messages the same way a person would, except it skips the 10 minutes of confused clicking before finding the right screen.
The credential wall
People talk about AI agents automating knowledge work, and they’re probably right about where this is heading. But the bottleneck in 2026 is not reasoning capability. It’s the 30 minutes of admin console clicking that happens before any agent can do anything useful.
OpenClaw is good once it has Gmail access. But getting it that access meant navigating one of the most confusing UIs in tech. As more people adopt AI agents — OpenClaw, n8n, LangChain, whatever appears next — they all hit this same wall. Every agent needs API keys, OAuth tokens, webhook URLs, service account credentials. And getting those means logging into web consoles that were built for infrequent use by people who supposedly already know where everything is.
A browser agent that operates these UIs is not a nice-to-have for anyone running agentic AI tools. It’s infrastructure.