Why this test exists
A practical cross-agent website test planner
Results from different agents are only comparable when the goal, starting state, permissions, environment, assertions, and readback stay consistent. Ad hoc prompts create anecdotes; a shared GoalContract creates a reproducible evaluation plan.
Cross-Agent Replay exports that plan for Grok, Claude, ChatGPT Work, and Gemini. Every provider row starts as not_run because this public tool does not open provider accounts or claim results it did not observe. This release does not edit or import provider receipts; teams preserve separately authorized results alongside the exported plan.