Gemini 3.1 Pro Just Dropped. Dassi Already Runs It.
Google shipped Gemini 3.1 Pro on February 19th, and the benchmark numbers grabbed me before the marketing copy did. ARC-AGI-2 score: 77.1%, more than doubling Gemini 3 Pro’s 31.1%. That is not an incremental update. That’s a different model wearing the same name.
Dassi supports it today. If you have a Google AI Studio API key, you can plug it in and start running browser tasks with it right now.
What actually changed under the hood
The standout feature is a three-tier thinking system. Previous Gemini models gave you a binary toggle between low and high compute modes, which always felt like choosing between a quick guess and an expensive meditation with nothing practical in between and no middle ground for the kind of moderately complex tasks that fill most of an actual workday. Gemini 3.1 Pro adds a medium tier that sits right where most real work happens.
Google’s thinking_level parameter controls this. Low for autocomplete-style speed. Medium for code review and moderately complex analysis. High for deep debugging where you want the model to slow down and really chew on the problem. Default is high, which is generous if you’re working on something hard and wasteful if you’re not.
Context window stays at 1M tokens. But the output ceiling jumped to 65,536 tokens, and that matters more than it sounds because Gemini 3 Pro had an infuriating habit of truncating code generation around 21,000 tokens. You would get three-quarters of a response and then silence. So that ceiling is gone.
Benchmarks worth looking at
GPQA Diamond for scientific knowledge: 94.3%. SWE-Bench Verified for agentic coding: 80.6%, putting it a hair below Claude Opus 4.6’s 80.8%. MMMLU for multimodal understanding: 92.6%. These are Google’s published numbers from the model card.
Speed tells a stranger story though. Artificial Analysis clocked output at 104.7 tokens per second, solidly above the 71.6 median for comparable frontier models. But time to first token sits at nearly 34 seconds, when the median for similar models is about 1.2. So you stare at a blank screen for half a minute, then text floods in all at once. For quick back-and-forth browser interactions like filling forms or clicking through navigation sequences, that initial wait is a real obstacle. For research-heavy analysis where you’re feeding in large pages and expecting synthesis, probably fine.
I do not have verified data on how Gemini 3.1 Pro performs specifically on browser automation benchmarks, so I won’t make claims there. Run it yourself and see.
Plugging it into dassi
If you’ve used dassi with Claude or GPT before, setup is identical. Grab a key from Google AI Studio, paste it into dassi’s settings, select Gemini 3.1 Pro. Your key stays in your browser. Nothing routes through our servers.
And this is why BYOK matters as a principle. When a new model drops, you do not wait for anyone to negotiate a partnership or roll out an integration. You switch. Or switch back if it doesn’t fit your tasks. We wrote about why model flexibility matters for browser agents a few weeks ago, and every major release keeps reinforcing that argument.
The cost trap
Pricing matches Gemini 3 Pro exactly. $2 per million input tokens under 200K context, $12 per million output tokens. So it is effectively a free performance upgrade on paper.
But the thinking tiers create a trap for the inattentive. A complex request on high mode can burn through 30,000+ thinking tokens at the output rate, costing around $0.36 per request. That same request on low mode uses about 1,000 thinking tokens for $0.012. Thirty times cheaper. For browser sessions where you’re sending dozens of requests, that gap compounds damn fast. Pay attention to which tier you’re using.
Still in preview
Google has not moved this to general availability yet.
Try it and compare
Best way to find out if Gemini 3.1 Pro fits your workflow is to just swap it in and run your usual browser tasks. Because nobody can tell you which LLM works best for the specific stuff you do better than a few minutes of actually testing it yourself. The 1M context window is genuinely useful for long-page analysis, and the three-tier thinking gives you cost control that most frontier models still lack. Whether the 34-second cold start bothers you depends entirely on what you’re doing with it.