Comparable replay plan

Planning artifact · Free browser tool

Cross-Agent Replay

Export a provider-neutral replay packet before testing. This release creates the shared GoalContract and four explicit not-run evidence rows; it does not model or import provider-specific execution receipts.

1 · Define and run

GoalContract input

Start with the included example, then replace it with one narrow customer goal and its real authority boundary.

Provider-neutral export

Grok · not_runClaude · not_runChatGPT Work · not_runGemini · not_run

This workbench exports comparable rows only. It does not contact a provider, open a browser, consume provider credits, or publish results.

Authority, assertions, and safe stop

Selecting a sandbox records a declaration in the contract; a later executor must still verify written authorization, isolation, fixture reset, and revocation. This public tool does not execute the sandbox run.

By running this tool, you confirm the submitted inputs are within your authorized scope.

2 · Inspect and export

Evidence bundle

Your structured result will appear here

Review the example GoalContract, adjust its boundaries, and run the tool. No action is taken beyond the boundary shown on this page.

Why this test exists

A practical cross-agent website test planner

Results from different agents are only comparable when the goal, starting state, permissions, environment, assertions, and readback stay consistent. Ad hoc prompts create anecdotes; a shared GoalContract creates a reproducible evaluation plan.

Cross-Agent Replay exports that plan for Grok, Claude, ChatGPT Work, and Gemini. Every provider row starts as not_run because this public tool does not open provider accounts or claim results it did not observe. This release does not edit or import provider receipts; teams preserve separately authorized results alongside the exported plan.

Repeatable workflow

From a bounded goal to an inspectable receipt

  1. 1

    Freeze the common contract

    Define one goal, start URL, action boundary, success assertions, safe stop, and authoritative readback.

  2. 2

    Export provider rows

    Create a replay plan with one comparable row per provider, all explicitly marked not_run.

  3. 3

    Run in authorized harnesses

    Hold the contract and environment stable while each provider executes under its own approved account and controls.

  4. 4

    Attach receipts before comparison

    After separate authorized runs, preserve evidence coverage and authoritative outcomes beside this plan, then compare compatible receipts instead of prose impressions.

Read the evidence precisely

Four interpretation rules

not_run is a real state

An unexecuted provider row is neither a failure nor a zero and must not affect a success denominator.

Harness differences matter

Browser versions, account state, model configuration, prompts, and tool availability belong with each result.

Same goal, same stop

Changing the success bar or allowing one provider to cross a consequential boundary makes the comparison invalid.

Permission precedes publication

Only permissioned, methodologically compatible results should feed public comparisons or leaderboards.

Included in this tool

Observable checks and exports

  • One content-derived GoalContract revision shared across four provider result slots
  • Four explicit not-run evidence rows, one for each named provider surface
  • Shared safe-stop, success-assertion, and authoritative-readback requirements
  • Portable JSON for permissioned manual collection or future integrations

Keep outside the claim

Known limitations

  • No Grok, Claude, ChatGPT Work, or Gemini session is started.
  • Provider availability, plans, regions, and harness versions must be recorded at actual run time.
  • Leaderboard publication requires consent, comparable coverage, methodology metadata, and verified receipts.

Frequently asked questions

Does this automatically run all four agents?

No. It creates the shared contract and four explicit not-run records. Automated provider execution would require separate authorized integrations and controlled accounts.

Why use one contract for every provider?

Keeping the goal, starting state, boundaries, assertions, and fixtures fixed makes differences easier to attribute to the tested surface instead of a changed task.

Can the results go on AgentReady leaderboards?

Only after permissioned execution produces comparable, verified receipts and the owner explicitly consents to publication under the stated methodology.

Continue the investigation

Related tools and field research

Review the opt-in leaderboard model

Public scan projection

Agent Access Matrix

Compare robots policy, simulated agent user-agent responses, server-rendered content, browser stability, CAPTCHA, and security evidence in one public-site preflight.

Open tool

Local JSON analysis

Lighthouse Agentic Importer

Import a Lighthouse JSON report locally, preserve official audit IDs and display values, and turn observed failures into an agent-journey rerun checklist.

Open tool

Planning artifact

Journey Contract Builder

Define the goal, starting state, allowed and prohibited actions, safe stop, assertions, and authoritative readback before an agent touches a site.

Open tool

Measured comparisons

Use receipts—not anecdotes—in a leaderboard

AgentReady's public leaderboard model requires owner opt-in, category fit, compatible scanner versions, observation windows, denominators, and evidence coverage. A tool export is an input to that process, not automatic publication.

View leaderboards

Need the whole public-site baseline?

Run the free AgentReady scan for discovery, semantics, browser compatibility, public forms, safety signals, and evidence-backed fixes.

Scan a public URL