Benchmark blueprint
Grok Bot vs. Claude in Chrome vs. ChatGPT Work vs. Gemini: A 10-Task Website Test
A fair ten-task blueprint for comparing Grok Bot, Claude in Chrome, ChatGPT Work, and Gemini without presenting proposed tests as observed results.
By AgentReady Editorial Team. Technical review: AgentReady Engineering.
Current official surface: Grok Bot runs on its own computer
xAI introduced Grok Bot in August 2026 as an early-beta, always-on teammate with a computer of its own. The official launch says Bots can sign into tools, work across applications and websites, continue after the user steps away, and handle products that lack a clean API or MCP surface. xAI subsequently expanded plan availability. Those facts justify testing Grok Bot as a persistent computer-using product, not merely sending a webpage to a Grok language model and calling the result a browser benchmark.12
The launch pages do not define a universal website conformance test, publish a site-owner crawler identity, or promise that every workflow will complete. A benchmark must record the Grok Bot application version, account plan, region, starting computer state, installed integrations, saved routine state, and whether the Bot has previously seen the site. Persistence and learned routines can be product benefits, but they are confounders in a cold-start comparison. Use a fresh Bot or report prior exposure explicitly.12
Current official surface: Claude spans Chrome and Cowork's built-in browser
Anthropic made Claude in Chrome generally available on paid plans and describes it as able to view the current page, read and type text, click, navigate, fill forms, and use existing logins. Anthropic also says a safety classifier checks autonomous actions against the user's request, with prompt-injection defenses remaining part of the product design. Separately, Claude Cowork now has a built-in browser inside the desktop application. Those are two browser surfaces with different starting state: an existing Chrome session may already contain cookies, extensions, and open tabs, while a built-in browser has its own environment.34
Therefore the comparison row cannot simply say Claude. It must say Claude in Chrome or Claude Cowork built-in browser, along with approval mode, account plan, selected model when configurable, and authentication state. If a site works only because the local Chrome profile is already signed in, that is legitimate evidence for that surface but not proof that a fresh remote browser works. If the safety layer refuses or pauses on a risky instruction, record a safety outcome rather than treating it as a navigation defect.34
Current official surface: ChatGPT Agent has transitioned to ChatGPT Work
The current OpenAI help surface says the former ChatGPT Agent mode is no longer available and directs users to ChatGPT Work and its cloud browser. Accordingly, this article's search-oriented title retains ChatGPT, but the actual benchmark target is ChatGPT Work cloud browser. That cloud browser runs on a separate remote computer, can read pages, click, enter information, and handle supported public and signed-in tasks. It pauses for user input, sign-in, or confirmation, and the secure sign-in flow sends credentials to the remote browser rather than exposing them to the model.56
OpenAI also documents Web Bot Auth for cloud-browser traffic, allowing site operators to verify signed requests at supporting edge providers or with published keys. In the ChatGPT desktop app's built-in browser, site tools use the proposed WebMCP standard so a participating website can expose structured operations that ChatGPT may use instead of only clicking and typing. These are distinct surfaces. A cloud-browser run and a desktop built-in-browser site-tool run must occupy separate benchmark rows, and an older ChatGPT Agent result should be labeled historical rather than silently mixed into current data.78
Current official surface: Gemini in Chrome uses auto browse
Google's current Gemini in Chrome help documentation describes auto browse for multi-step tasks such as comparing products, adding items to a cart, finding travel accommodations, making reservations, and scheduling appointments. The product can ask the user to review a plan, confirm sensitive steps, take over, or grant sign-in permission. Depending on the flow it may use local Chrome state or a remote browser, and Google documents safeguards and user responsibility because prompt injection and unintended actions remain possible.9
Availability, plan, account type, region, device, and rollout state are part of the evidence. A test that one evaluator can run in a local signed-in browser may be unavailable to another evaluator or may resume in a remote session with a different state. Record which path was actually used, whether Password Manager assisted sign-in, and every take-over or confirmation event. Do not infer that a Gemini API model or a Google search result reproduces Gemini in Chrome auto browse.9
Current AgentReady observation: public prerequisites, not these four provider runs
As reviewed on August 30, 2026, AgentReady's public diagnostic checks discovery policy, semantic HTML, accessible controls, content structure, browser stability, public safety signals, legal surfaces, and conditional API, MCP, repository, and commerce evidence. Its enhanced browser can attempt bounded navigation, fill a safe public form without submitting it, find documentation or contact information, and inspect a declared sandbox. It blocks non-GET and non-HEAD requests. AgentReady can produce a Fix Pack and verify stable findings after deployment.10
AgentReady does not currently sign into Grok Bot, Claude, ChatGPT Work, or Gemini accounts and run this cross-provider matrix. Its public runner uses a fresh anonymous browser context and does not receive credentials for the scanned customer's product; it does not currently make purchases or create accounts. It aborts observed non-GET and non-HEAD requests, although that cannot guarantee safety on a server incorrectly implemented to mutate on GET. A strong AgentReady score is useful prerequisite evidence, not a completed result for this benchmark. The distinction is explained in agent readiness score versus journey success. Any future comparison must preserve the public scan and provider journey run as separate evidence ledgers.10
Benchmark blueprint: choose a common task with provider-specific safe stops
Begin with a task that every available surface is permitted and reasonably expected to attempt. A public discovery contract could say: starting from the homepage, find the least expensive plan that includes feature X, report the price and two constraints, and cite the canonical source URL and as-of date. A public interaction contract could say: locate the demo form, identify every required field, populate synthetic values if permitted, and stop before submission. These tasks test retrieval, navigation, semantics, state recognition, and restraint without requiring money or customer data.1112
For post-login work, create a disposable test tenant and follow how to test AI agents after login safely. The shared goal might be to locate a synthetic invoice and report its due date without editing, downloading, messaging, or paying. If one provider's current policy prohibits a step that another supports, set a provider-specific safe stop and mark the excluded step not applicable. Never prompt a system to bypass CAPTCHA, evade the WAF, defeat an approval, or violate its own safety controls merely to keep the rows symmetrical.59
- Freeze one versioned journey contract: Name the goal, exact task, start URL, fixture, allowed and prohibited actions, safe stop, positive assertions, negative assertions, timeout, retry rule, and observation window. Publish the contract before examining provider outcomes.
- Create a clean environment for each surface: Use fresh product sessions where possible, equivalent synthetic accounts, identical locale and viewport, and no prior exposure to the target. Record unavoidable differences such as local Chrome state, remote browser state, plan availability, or learned routines.
- Run controls before agents: Verify that an ordinary browser can reach the page, the fixture exists, expected content is live, and the site is not generally broken. Run AgentReady's public diagnostic to preserve deterministic prerequisites and coverage, but keep that score outside the provider outcome column.
- Capture the first attempt and every intervention: Preserve the initial plan, first blocking or ambiguous state, relevant rendered evidence, final URL, approval prompts, takeovers, retries, and sanitized receipts. A rescued task is human-assisted, not an autonomous first-attempt success.
- Verify assertions outside the agent: Check the cited URL, extracted facts, audit logs, fixture state, and absence of prohibited side effects. The agent's summary is not proof. Mark completed only when every required assertion passes.
Use evidence states, not a winner-take-all score
Report each run as completed, partial, blocked by site, refused by provider policy, unsafe, unavailable, or unobservable. Add first-attempt completion, intervention count, retries, step count, citation correctness, assertion coverage, approval correctness, and time as secondary fields. A refusal at a prohibited purchase is not equivalent to a login page hidden from accessibility APIs. A rollout restriction is not a model failure. A WAF block should link to the access evidence described in robots.txt and WAF testing for AI agents.59
Avoid a universal model leaderboard. Results belong to a product surface, version, task, account state, and date. If aggregation is useful, publish cohorts with disclosed denominators, minimum sample sizes, missing-data treatment, ties, and methodology changes. AgentReady's proposed leaderboards provide an opt-in publication boundary; the Agentic Customer Journey Index demonstrates dated task contracts, safe stops, source records, coverage, and correction routes without claiming an unconsented numeric winner. This blueprint proposes future cross-provider rows and does not populate them.
Test both the visual interface and structured site tools
A fair website evaluation should preserve a visual or semantic browser path because not every agent supports site tools and not every website exposes them. Stable headings, accessible names, native controls, explicit state, useful errors, and durable receipts remain the baseline. Google's agent-ready toolkit frames deterministic browser audits as a way to improve these conditions. Those audits help explain failures but do not replace the provider journey result.131112
Where current surfaces support WebMCP or site tools, add a separate structured-tool run. OpenAI says site tools can work with the current page and signed-in session and can expose read or change operations, subject to browser permissions and confirmation requirements. The WebMCP website-tools guide explains the implementation layer. Do not mix a tool-assisted success with a click-only result; report whether the agent discovered the right tool, supplied valid inputs, respected authorization, handled the output, and verified the visible page state.8
Available replay plan and a future cross-harness runner
The available Cross-Agent Replay creates one versioned GoalContract and four explicit not-run evidence rows for Grok, Claude, ChatGPT Work, and Gemini. The shared contract freezes the scope, boundaries, safe stop, assertions, and authoritative-readback requirement before testing begins. This release does not edit or import provider result rows, contact a provider, open a browser, consume credits, collect passwords, or publish a ranking. A full Cross-Harness Journey Runner remains proposed and would attach separately authorized, privacy-safe provider receipts with intervention and first-failure fields.
A proposed Harness Change Log would monitor official provider pages and mark an old comparison stale when a product transitions, as ChatGPT Agent has transitioned to ChatGPT Work, or when a new browser or structured-tool surface appears. That prevents an SEO page from quietly comparing discontinued and current products as peers. Use the available replay plan and AgentReady public scan for prerequisites, follow the provider-specific safety documentation, and publish only evidence actually collected under consent. Also review Is your site usable by Grok Bot? and Claude in Chrome website readiness for the product-specific interpretation behind the shared matrix.1359
Adjacent AgentReady tools
Available tools turn the article into a bounded check or reusable contract. Proposed tools remain roadmap candidates and are not claimed as live.
Conclusion
The meaningful comparison is not Grok model versus Claude model versus ChatGPT model versus Gemini model. It is Grok Bot on its own computer, a named Claude browser surface, the current ChatGPT Work cloud browser or desktop site-tool surface, and Gemini in Chrome auto browse, each attempting the same versioned journey under documented conditions. The former ChatGPT Agent label must remain historical; current tests should name ChatGPT Work. This article supplies the blueprint, not results: establish controls, freeze the task and safe stop, use clean fixtures, record provider applicability, capture the first attempt, verify state outside the agent, and publish evidence states with dates and denominators. AgentReady can provide the deterministic public baseline and provider-neutral replay plan today. A future cross-harness runner can connect those artifacts to provider journey evidence without manufacturing a universal winner.
Compare evidence across real sites
Use the research index for public, task-specific observations. Leaderboards remain methodology-controlled and require owner opt-in before numeric ranking or named improvement claims.
Sources
Each source shows its individual verification date. Recheck current versions before relying on time-sensitive requirements.
- Introducing Grok Bot — xAI; checked August 30, 2026
- Grok Bot is now available on more plans — xAI; checked August 30, 2026
- Claude in Chrome is now generally available — Anthropic; checked August 30, 2026
- Cowork gets a built-in browser — Anthropic; checked August 30, 2026
- Using cloud browser in ChatGPT — OpenAI Help Center; checked August 30, 2026
- ChatGPT agent — OpenAI Help Center; checked August 30, 2026
- ChatGPT Work cloud browser allowlisting — OpenAI Help Center; checked August 30, 2026
- Using site tools in the ChatGPT desktop app — OpenAI Help Center; checked August 30, 2026
- Ask Gemini in Chrome to complete tasks with auto browse — Google Gemini Help; checked August 30, 2026
- AgentReady scoring methodology and limitations — AgentReady; checked August 30, 2026
- Web Content Accessibility Guidelines 2.2 — W3C; checked July 13, 2026
- HTML Living Standard — WHATWG; checked July 13, 2026
- An agent-ready toolkit for the web — Chrome for Developers; checked August 30, 2026
Related resources
Apply this to a real outcome
Use the goal-specific playbooks to turn this guide into a task contract for discovery, signup, booking, commerce, or product use.