Crawler and browser access preflight

Public scan projection · Free browser tool

Agent Access Matrix

Run one existing AgentReady public scan and turn the relevant observations into an access matrix. The result keeps policy, edge behavior, and real-browser evidence separate so a passing robots.txt check cannot hide a WAF or rendering failure.

1 · Define and run

GoalContract input

Start with the included example, then replace it with one narrow customer goal and its real authority boundary.

Public baseline: this submits only the starting URL to AgentReady's existing public scan with browser checks enabled. It does not request enhanced account access or publish a report.

Authority, assertions, and safe stop

Selecting a sandbox records a declaration in the contract; a later executor must still verify written authorization, isolation, fixture reset, and revocation. This public tool does not execute the sandbox run.

By running this tool, you confirm the submitted inputs are within your authorized scope.

2 · Inspect and export

Evidence bundle

Your structured result will appear here

Review the example GoalContract, adjust its boundaries, and run the tool. No action is taken beyond the boundary shown on this page.

Why this test exists

A practical AI agent crawler access checker

A permissive robots.txt rule is not proof that an AI agent can retrieve or use a page. A CDN, web application firewall, consent layer, JavaScript shell, redirect chain, or bot-verification rule can produce a completely different result after the crawler policy is read.

The Agent Access Matrix projects one public AgentReady scan into separate access claims. That separation matters: published policy, ordinary HTTP retrieval, rendered browser access, and verified crawler identity answer different questions and require different evidence.

Repeatable workflow

From a bounded goal to an inspectable receipt

  1. 1

    Name the retrieval goal

    Describe the public page or discovery task, allowed origin, and safe stopping point.

  2. 2

    Run the public baseline

    AgentReady requests the submitted public URL and includes the existing browser-based checks.

  3. 3

    Inspect each access plane

    Review policy, edge behavior, retrieval, and browser evidence without collapsing unknowns into failures.

  4. 4

    Export and repeat

    Save the contract, run, evidence, and receipt as JSON, then compare it after an edge or policy change.

Read the evidence precisely

Four interpretation rules

Policy is intent

robots.txt communicates crawler preferences; it does not prove that the edge delivered the requested representation.

Retrieval is an observation

A successful public response supports reachability for that test path and time, not universal access.

Browser parity is separate

Rendered content may differ from the initial response because of JavaScript, consent, authentication, or bot controls.

Identity needs verification

A user-agent string alone is not crawler identity. Verified reverse DNS or provider-specific mechanisms must be checked separately.

Included in this tool

Observable checks and exports

  • One URL-only scan using AgentReady’s existing SSRF protections and public rate limits
  • Separate rows for declared crawler policy, edge response, rendered content, and browser behavior
  • Evidence-preserving pass, partial, fail, not-applicable, and unobservable states
  • Downloadable GoalContract, TestRun, evidence, and OutcomeReceipt JSON

Keep outside the claim

Known limitations

  • No reverse-DNS or provider-IP verification is performed.
  • A public preflight cannot observe authenticated pages or private account state.
  • The receipt describes this scan only; it is not a certification or indexing guarantee.

Frequently asked questions

Does this test whether Grok can crawl my website?

It tests the public conditions Grok and other agents may encounter, including robots policy and simulated user-agent responses. It does not authenticate a request as xAI’s crawler or guarantee Grok retrieval.

Why can robots.txt pass while agent access still fails?

robots.txt expresses a site policy. A CDN, WAF, CAPTCHA, JavaScript dependency, redirect, or browser error can still prevent retrieval or interaction after policy allows it.

Does running the matrix publish my scan?

No. The tool uses the public scan endpoint without opting into a public report. It returns evidence to the current browser session and does not create an indexable report.

Continue the investigation

Related tools and field research

Compare the leaderboard methodology

Local JSON analysis

Lighthouse Agentic Importer

Import a Lighthouse JSON report locally, preserve official audit IDs and display values, and turn observed failures into an agent-journey rerun checklist.

Open tool

Planning artifact

Journey Contract Builder

Define the goal, starting state, allowed and prohibited actions, safe stop, assertions, and authoritative readback before an agent touches a site.

Open tool

Public scan projection

Stripe Link Checkout Preflight

Check the public substrate an agent needs before checkout: product data, structured prices, price parity, policies, HTTPS, and discoverable interaction evidence.

Open tool

Measured comparisons

Use receipts—not anecdotes—in a leaderboard

AgentReady's public leaderboard model requires owner opt-in, category fit, compatible scanner versions, observation windows, denominators, and evidence coverage. A tool export is an input to that process, not automatic publication.

View leaderboards

Need the whole public-site baseline?

Run the free AgentReady scan for discovery, semantics, browser compatibility, public forms, safety signals, and evidence-backed fixes.

Scan a public URL