# Agent Visibility Score
## Methodology v1.0

**Owner:** The Future Collective  
**Version:** 1.0  
**Date:** 25 September 2026  
**Status:** Controlled-pilot specification; not yet externally validated  
**Initial market:** UK retail and consumer commerce  
**Review point:** After five design-partner assessments and the first comparable index wave

---

## Executive decision

The Agent Visibility Score (AVS) measures how well a defined organisation and its commercial proposition can be accessed, interpreted, trusted, found, recommended and acted on by a specified panel of AI systems and agent tests, in a specified market and time window.

The score has six dimensions:

1. **Access**
2. **Read**
3. **Trust**
4. **Discover**
5. **Recommend**
6. **Transact**

Each dimension receives an equally weighted score from 0 to 100. The headline score is the arithmetic mean of the six dimension scores. Equal weights are a transparent starting convention. They are not a claim that the dimensions have equal commercial impact, and they must be reviewed after pilot evidence has been collected.

Every reported score must be accompanied by its six dimension scores, evidence coverage, confidence grade, sample definition, systems tested and observation dates. A score without those details is incomplete.

The score measures **observed performance within the stated test scope**. It does not predict sales, certify that a business is ready for every agent, or establish what an organisation can do through private systems that were not authorised for testing.

## 1. Purpose and governing question

### 1.1 Governing question

> **Can an AI system acting for a customer reach this organisation, understand its offer, establish whether it is trustworthy and suitable, recommend it when relevant, and progress an authorised task towards a safe transaction?**

AVS is designed to make that question measurable for a particular organisation, market, category and observation period. The first release is UK retail. Later sector modules may adapt the product and task samples, but the six dimensions and evidence rules remain stable unless a controlled method revision changes them.

### 1.2 What the score is for

- A repeatable baseline for one organisation or UK retail proposition.
- A like-for-like comparison within a defined cohort and common test panel.
- A starting point for an evidence-led diagnostic and remediation plan.
- Longitudinal monitoring of changes in public machine access, product truth and agent outcomes.
- A consistent structure for the Agent Visibility Snapshot, Diagnostic, Benchmark and Monitor.

### 1.3 What the score is not

- A prediction of revenue, conversion or future market share.
- A score of general search engine optimisation or generative-engine optimisation.
- A guarantee that an agent will recommend or buy from a business.
- A security assessment, legal opinion or assurance of regulatory compliance.
- A measure of private feeds, APIs, credentials, internal governance or contracts unless these are supplied and authorised for a separate diagnostic.
- A permanent judgement. Agent products, models, crawlers, interfaces and policies change.

### 1.4 Measurement unit

The unit of analysis is a **specific commercial proposition**, not a corporate group in the abstract. Before fieldwork, record:

- trading name and, where publicly available, legal seller;
- parent company and marketplace or concession relationships;
- canonical UK domain, relevant subdomains and app dependence;
- category, currency, fulfilment scope and consumer audience;
- whether checkout remains with the named retailer or transfers elsewhere;
- test period, systems, locale and any required location setting.

Separate propositions must be scored separately where seller, catalogue, checkout, fulfilment or recourse materially changes.

## 2. Relationship to the existing research standard

The Future Collective’s Wave 4 evidence framework defines a public Agentic Commerce Readiness Index around four journey gates: **Find, Understand, Act and Resolve**. It also defines a broader enterprise diagnostic with internal evidence. That work remains an important companion standard.

AVS adds a six-dimension commercial measurement layer. It uses shared evidence where appropriate, while separating the questions each product answers:

| Measure | Primary question | Evidence boundary |
|---|---|---|
| **Agent Visibility Score** | How does this defined proposition perform across access, machine interpretation, trust, discovery, recommendation and transaction tests? | Outside-in evidence and named, controlled agent tests; authorised client evidence may be added only in a separately labelled diagnostic. |
| **Agentic Commerce Readiness Index** | How ready is the observable market-facing journey, including resolution after a transaction? | Public and genuinely available consumer surfaces, across Find, Understand, Act and Resolve. |
| **Enterprise Agentic Commerce Readiness Diagnostic** | What internal capabilities, controls, ownership and investment decisions explain the outside-in result? | Outside-in evidence plus authorised internal documents, interviews, systems, logs and controls. |

Do not silently convert the previous four-gate index into the six-dimension AVS, combine their weights, or present their totals as directly comparable. If both measures are published for the same retailer, display their full names, questions and coverage separately. For a league table, select one headline measure and freeze its method before scoring begins.

The conceptual crosswalk is:

| Existing journey gate | AVS dimensions that inform it | Limit |
|---|---|---|
| Find | Access + Discover | AVS tests the named systems and routes in its panel. |
| Understand | Read + Trust | AVS tests sampled facts and sources, not every internal data control. |
| Act | Recommend + Transact | AVS tests selected missions and a safe pre-purchase boundary. |
| Resolve | No full AVS equivalent | AVS indicator X4 checks whether recovery and recourse routes are visible. It does not prove that a completed order can be cancelled, returned or resolved end to end. |

This crosswalk explains the relationship; it is not a conversion formula.

AVS preserves the prior framework’s evidence states, no-silent-zero rule, category relevance gates, fieldwork freeze, independent review, claims discipline and versioned evidence record.

## 3. The six dimensions

The existing four-part site language, **Access, Read, Trust, Transact**, remains intact. **Discover** and **Recommend** make the market outcome visible alongside the capabilities that support it.

| Dimension | Weight | Question | Boundary |
|---|---:|---|---|
| **Access** | 1/6 | Can the relevant systems reach the intended public information or permitted action route? | Measures named paths and systems. A crawler policy is not proof that all AI systems are blocked. |
| **Read** | 1/6 | Can machine-readable and rendered information identify the right product, variant, attributes and offer? | Measures tested fields and consistency, not the mere presence of a particular schema type. |
| **Trust** | 1/6 | Can an agent establish who is selling, which source is authoritative, whether material facts agree and where policies or claims come from? | Evidence-based verifiability, not a subjective judgement of brand reputation. |
| **Discover** | 1/6 | Do named systems find the relevant organisation, product or authoritative source for a defined query? | A dated observation of a fixed panel and prompt set, not universal visibility. |
| **Recommend** | 1/6 | Does the proposition appear accurately and appropriately in defined shopping missions? | Measures observed recommendation outcomes, not actual consumer preference or sales. |
| **Transact** | 1/6 | Can a controlled, authorised agent progress a task through product selection, basket and a safe handoff or review boundary? | No purchase is made. Private or unavailable routes remain unknown unless authorised evidence is provided. |

Dimensions have equal weight, calculated as an exact 1/6 each (approximately 16.67% when displayed). Each dimension is calculated first; indicator counts do not change the dimension weight. This prevents a dimension with more tests from dominating the headline score.

## 4. Criterion dictionary

The following indicators form the minimum common v1 instrument. A category module can add criteria, but must not remove a relevant core criterion or change its meaning without a method version change.

### 4.1 Access

| ID | Indicator | Minimum evidence and scoring basis |
|---|---|---|
| A1 | Search-crawler policy compatibility | Capture robots.txt and other public instructions. Classify each named crawler by its documented purpose. Score the share of eligible tested paths that the applicable search/indexing crawler policy permits; show intentional restrictions separately as policy choices, not technical defects. Keep training and user-directed fetcher rules separate. |
| A2 | Public endpoint reachability | Test the frozen homepage, category, product and policy URLs with a normal, identified, unauthenticated client. Record status, redirect chain and semantic page result. A challenge or deny is a result; do not bypass it. |
| A3 | Content availability through the tested route | Check whether the relevant page content is returned or rendered in the named, permitted retrieval or consumer-system route. A generic user-agent string is a simulation, not a genuine provider test. |
| A4 | Access stability | Repeat valid access checks in two fieldwork windows and two UK network vantages. Score successful, semantically usable responses; retain transient failures and retries. |

Robots rules express crawler preferences. They are not access authorisation and do not establish whether a user-directed fetch, browser agent, shopping assistant or commerce protocol can transact. Score a restriction only for the exact relevant crawler and path tested. Keep the reason visible beside the score. If no applicable policy can be retrieved or interpreted, mark A1 unknown and reduce coverage; do not assume either permission or denial.

### 4.2 Read

| ID | Indicator | Minimum evidence and scoring basis |
|---|---|---|
| R1 | Entity and variant identity | For each sampled product, verify that product, brand, seller, variant and stable identifiers can be distinguished. Record GTIN, SKU, model or other relevant IDs where available; absence is not automatically a failure when the identifier is not applicable. |
| R2 | Decision-relevant field coverage | Check category-relevant fields such as price, currency, availability, size, colour, pack, ingredients, compatibility, dimensions or material. Score the proportion of applicable fields present and interpretable. |
| R3 | Machine extraction | Extract relevant content from rendered text and available structured sources such as JSON-LD, feeds or documented interfaces. Record which source is being tested; markup by itself does not prove that a consumer agent uses it. |
| R4 | Visible-to-machine parity | Compare extracted name, variant, price, availability and material attributes with the visible page at the same time and location state. Score the proportion of tested facts that agree. |

### 4.3 Trust

| ID | Indicator | Minimum evidence and scoring basis |
|---|---|---|
| T1 | Seller and source authority | Can the tested system identify the seller, canonical product source and any marketplace or fulfilment relationship? Verify against first-party pages and public records where relevant. |
| T2 | Offer truth | Compare price, currency, promotion conditions, unit price where relevant, stock and mandatory costs across visible page, extracted data, agent answer and basket state where available. |
| T3 | Policy clarity | Test whether delivery, returns, eligibility, subscription or other material conditions can be found and attributed before a commitment point. Score clarity and consistency, not the generosity of the policy. |
| T4 | Claim and evidence traceability | For material suitability, safety, provenance or performance statements in the sample, record whether the claim has an identifiable first-party or authoritative source. Do not award points for unsupported claims repeated across channels. |

Trust does not measure whether a brand is reputable or whether its products are objectively good. It measures whether the relevant facts and responsibilities can be checked within the defined test.

### 4.4 Discover

| ID | Indicator | Minimum evidence and scoring basis |
|---|---|---|
| D1 | Correct entity retrieval | In fixed, brand- or product-specific prompts, does the system return the correct UK proposition or product rather than a similarly named entity? |
| D2 | Authoritative source retrieval | Where the system supplies links or citations, does it surface the correct first-party product, policy or retailer source? Record systems without citations as not applicable to this indicator, not as a failure. |
| D3 | Offer and market accuracy | When a relevant entity is found, are the UK market, current product or variant, and material offer details correctly scoped? Score factual accuracy against the captured source evidence. |

### 4.5 Recommend

| ID | Indicator | Minimum evidence and scoring basis |
|---|---|---|
| M1 | Recommendation inclusion rate | Share of valid open-ended shopping responses in which the proposition or a relevant product appears in the first five recommendations. Duplicate mentions in one response count once. |
| M2 | Recommendation position | For appearances in the first five, score position using 100 for first, 80 for second, 60 for third, 40 for fourth and 20 for fifth; no appearance scores zero for that response. Report the raw position distribution as well. |
| M3 | Constraint and fact fit | Share of surfaced recommendations that satisfy the prompt’s explicit constraints and have no material factual error in the fields checked. A wrong size, price, availability, seller or material constraint counts as a mismatch. |

**Share of Agent** is reported as a companion cohort metric, not as a substitute for the AVS. For each category mission, divide the number of valid top-five merchant/product slots occupied by the target proposition by all valid merchant/product slots returned across the same systems, prompts and runs. Deduplicate a proposition within a response. Report the denominator, panel, prompt set and period. Exclude or separately label sponsored placements.

### 4.6 Transact

| ID | Indicator | Minimum evidence and scoring basis |
|---|---|---|
| X1 | Correct selection | Can the test agent select the intended product, variant and quantity from the defined task without silent substitution? |
| X2 | Basket and total integrity | Does the basket preserve product identity, quantity, price and any visible fulfilment cost or condition at the point tested? |
| X3 | Safe progression and handoff | Can the agent progress to the authorised review or checkout handoff while preserving the user’s stated constraints and clearly stopping before an unapproved commitment? |
| X4 | Recovery and recourse visibility | Can the test identify how to amend or remove the basket and where cancellation, return or human-support information sits? This is a pre-purchase visibility test, not proof that a completed order can be resolved end to end. |

The public v1 transaction test stops before payment or order placement. A live transaction, post-purchase action or private protocol test requires explicit retailer authorisation and an agreed sandbox or test account. A documented API or partnership is recorded as a declared capability until a valid test demonstrates it.

## 5. Sampling and test protocol

### 5.1 Standard UK retail assessment

For one retailer or brand, freeze the sample before scoring:

- **Three product detail pages:** one routine/high-volume product, one materially variant-rich product and one constraint-sensitive product.
- **Five supporting pages:** UK homepage, one category page, delivery/fulfilment information, returns information and seller/contact or equivalent accountability page.
- **Four consumer-facing AI system slots:** initial panel candidates are ChatGPT with web search, Google Gemini or the relevant Google AI shopping/search surface, Microsoft Copilot and Perplexity. Record the exact product surface and model/version where available. A provider slot can only be changed before a wave is frozen, with the panel change documented.
- **Twelve prompts:** six discovery prompts and six recommendation missions, fixed before fieldwork and tailored only through pre-registered category variables.
- **Three independent runs per prompt per system:** 12 prompts × 3 runs × 4 systems = 144 response observations per assessed proposition. The minimum for a comparable score is three available systems and the same prompt/run structure across the cohort.
- **Access checks:** the eight URLs above in two fieldwork windows at least 24 hours apart and from two UK network vantages. Use the minimum requests needed, conservative concurrency and any published rate instructions.
- **Transaction checks:** three safe, pre-purchase tasks, each attempted twice on the same declared agent test harness or genuine supported consumer surface. If no valid transaction-capable test surface is available, mark the dimension unknown; do not substitute a conventional crawler test and call it agentic transaction evidence.

The standard sample supports a comparable outside-in score. A paid diagnostic should expand product sampling to at least twelve products across relevant categories and variants, then report the standard comparable score separately from the expanded findings. The expanded sample does not silently replace the comparable sample.

### 5.2 Prompt design

Every prompt is versioned, dated and stored exactly as run. Prompts must be neutral, concise and shopper-led. They must not name the target brand in open recommendation missions. Do not rewrite a prompt after seeing a result.

#### Discovery templates

1. “Find the UK website for [retailer] and show me where I can buy [product/category].”
2. “I am looking for [brand/product/model] in the UK. Find the official product page and tell me which version is available.”
3. “Where can I buy [product/identifier] from [retailer] in the UK?”
4. “Find [retailer]’s delivery information for online orders in the UK.”
5. “Find the UK returns policy for [retailer] and link to the source.”
6. “Find a [category] from [retailer] that matches [fixed, verifiable attribute]. Show the source page.”

#### Recommendation mission templates

1. “I need [category] for [use case], under £[budget], available from a UK retailer. Give me up to five suitable options and explain the fit.”
2. “Recommend a [category] that meets [constraint A] and [constraint B]. I am buying in the UK and need a current price and availability source.”
3. “I need a [product] suitable for [specific mission], with [attribute] and no more than £[budget]. Which options fit?”
4. “Compare up to five [category] options for [need]. Include the source, seller, price and any important limitation.”
5. “Find a [category] option for [household/user/context] that avoids [constraint] and is available for UK delivery.”
6. “I want to buy [category] this week. Recommend suitable options with current UK availability, total price if shown, and direct source links.”

Each prompt must have a corresponding expected-fact sheet: eligible category, explicit constraints, permissible interpretations, critical facts and evaluation rule. Prompts are adapted to each category before fieldwork freeze. Medical, age-restricted or other sensitive categories require specialist review and may be excluded from general recommendation scoring.

### 5.3 System controls

For each run, record:

- provider, consumer product, model/version if exposed and web/search mode;
- date/time in UTC, UK locale and location/postcode if required;
- account state, personalisation/memory settings and whether a fresh session was used;
- exact prompt and complete output, citations/links and visible sponsored placements;
- action steps, screenshots or screen recording where permitted;
- service error, refusal, outage, challenge or unexpected behaviour.

Use a fresh conversation for each run where possible. Randomise retailer and system order within the fieldwork schedule. Do not use private personal accounts or personalised purchase histories. If the product requires an account or postcode, use only an approved research account and documented test location.

### 5.4 Safe transaction protocol

The standard task begins from a publicly accessible product or category page. It may select an item, establish a variant and quantity, build a basket, inspect visible fulfilment choices and proceed to the order review boundary. It must not:

- place an order, submit payment or create a genuine customer commitment;
- enter a consumer’s personal, payment or authentication data;
- bypass a login, CAPTCHA, rate limit, consent boundary or technical control;
- test a private endpoint, API or account without written authorisation;
- continue after a step creates material purchase, privacy or operational risk.

Stop and record the boundary when a site requires login, identity verification, personal data, a purchase, an age/eligibility decision or a control bypass. An authorised sandbox may be used for a separate client diagnostic and must be labelled as such.

### 5.5 Robots and agent-purpose registry

Before every fieldwork wave, maintain a dated registry of relevant crawler and agent identities using primary provider documentation. At minimum, distinguish training crawlers, search/indexing crawlers, user-directed fetchers, browser/computer-use agents and commerce protocol clients. A provider may use different identities for different purposes.

Record the exact user-agent token, documented purpose, documentation date, matching robots group, applicable path rule and result. A simulated user-agent request can identify differential server treatment, but it cannot establish what the real provider experiences. Do not infer that a training opt-out prevents search inclusion or user-authorised retrieval.

## 6. Scoring rules

### 6.1 Indicator scoring

Use direct rates wherever a valid numerator and denominator exist:

**Indicator score = 100 × successful applicable observations ÷ valid applicable observations.**

For qualitative indicators that cannot be expressed as a direct rate, use the following anchored rubric and retain the written reason:

| Raw anchor | Score | Interpretation |
|---:|---:|---|
| 0 | 0 | Confirmed absent, materially wrong or consistently unusable in an appropriate completed test. |
| 1 | 25 | Fragmentary evidence; major gaps or failures. |
| 2 | 50 | Partial or inconsistent; the result works in some tested cases but material gaps remain. |
| 3 | 75 | Strong across most of the defined sample; remaining gaps are bounded. |
| 4 | 100 | Reliable and complete across the defined sample, repeated observations and relevant states. |

Do not award 100 merely because a feature exists in documentation. For a capability claim, distinguish **declared**, **observed**, **verified in a sandbox** and **verified in production**.

### 6.2 Dimension and total score

1. Calculate each indicator from valid observations.
2. Calculate the dimension score as the unweighted arithmetic mean of its applicable indicator scores.
3. Calculate the headline AVS as the unweighted arithmetic mean of the six dimension scores.
4. Round the headline score and dimension scores to whole numbers for display; retain full precision in the data.

Example: Access 75, Read 68, Trust 82, Discover 45, Recommend 34 and Transact 56 produce an AVS of 60. The report must show all six scores and coverage; “60” alone is not an adequate result.

### 6.3 Evidence coverage and confidence

Coverage is the share of planned, applicable evidence weight that was validly observed and scored. Unknown, unreachable or untested evidence reduces coverage; it is not silently assigned zero. Not-applicable criteria are excluded only when a written category relevance rule supports the decision. All six dimensions must have valid observations for an aggregate AVS.

| Publication status | Minimum rule | Permitted presentation |
|---|---|---|
| **Comparable** | At least 80% total evidence coverage; at least 60% coverage in each dimension; minimum three AI systems; full discovery/recommendation prompt protocol; transaction tasks attempted; second review complete. | Score and cohort comparison may be shown with scope, coverage and confidence. |
| **Provisional** | 60–79% total coverage; every dimension has some valid evidence; no material integrity dispute. | Show score as provisional and explain gaps. No league-table rank or “ready” label. |
| **Ungraded** | Below 60% coverage; any whole dimension unknown; insufficient panel observations; or unresolved material evidence/QA dispute. | Show the evidence profile and reason. Do not publish a headline score. |

Confidence is separate from performance:

- **A — High:** at least 85% coverage; two or more fieldwork windows; required UK access vantages; complete panel; second-review agreement of at least 95%; no material unresolved issue.
- **B — Good:** at least 80% coverage; repeat fieldwork and second review complete; bounded instability disclosed.
- **C — Limited:** 60–79% coverage or one important source remains unstable. Provisional presentation only.
- **U — Ungraded:** below 60%, a full dimension is unknown, or a critical integrity issue remains.

The score and confidence answer different questions. A low score can have high confidence; a high score can have low confidence.

### 6.4 Evidence states and missingness

| State | Meaning | Scoring treatment |
|---|---|---|
| Confirmed present | An appropriate completed test or reliable primary evidence supports the capability or fact. | Score the observed result. |
| Confirmed absent | An appropriate test was capable of detecting it and found it absent or consistently failing. | Score zero only for the relevant indicator. |
| Blocked | A policy, challenge or control prevented the observation. | May score zero for the tested access indicator; downstream indicators become unknown unless independently tested through another valid route. |
| Unreachable | Network/server failure or unresolved technical error prevented a valid observation. | Unknown for the score; report as friction. |
| Unknown | Available evidence cannot establish the position, or the route was outside authorisation. | Exclude from numerator and denominator; coverage falls. |
| Not applicable | A written category or business-model rule establishes that the criterion does not apply. | Exclude from numerator and denominator; reviewer approval required. |

Never convert “not found” into “does not exist” unless the test could reliably detect presence. Never convert a robots disallow into a conclusion about every agent or a protected resource.

### 6.5 No bands or certification in v1

V1 reports a number, profile, coverage and confidence. It does not use A–E grades, a “ready” badge, certification, public seal or pass mark. Those require pilot evidence, repeatability analysis, external review and a separately approved standard. This avoids implying that the initial equal-weight score has predictive or normative validation it does not yet have.

## 7. Evidence and data record

Every scored observation must be reproducible from a retained record. Store at least:

| Record group | Required fields |
|---|---|
| Method | method_version, panel_version, prompt_version, scoring configuration/checksum, test case ID |
| Entity and scope | retailer/proposition ID, seller, category, UK domain, product/variant, fieldwork window, currency and location state |
| System | provider, product surface, model/version if exposed, genuine or simulated identity, web/search mode, account state and locale |
| Request | timestamp_utc, URL or exact prompt, request profile, network vantage, retry/cooldown and deviation |
| Result | status, redirect chain, semantic page type, extracted fields, response/output, citations, action steps and completion state |
| Evidence | raw response or body where permitted, screenshot/DOM/payload, evidence URI, capture tool version, hash and retention classification |
| Score | indicator ID, applicability, evidence state, numerator, denominator, raw score, rationale, materiality and dimension assignment |
| Review | researcher, second reviewer, agreement, adjudication, challenge response and final decision |

Preserve original observations and method versions. Correcting a score creates a new version and a dated correction record; it does not overwrite the prior result. Evidence retention and access controls must respect provider terms, retailer permissions, privacy law and contractual commitments.

## 8. Fieldwork, quality assurance and challenge

### 8.1 Run sequence

1. Freeze the retailer entity, cohort, product sample, systems, prompts, criteria and method version before results are visible.
2. Capture the product and policy source pages, seller identity and relevant state.
3. Run access checks in two windows and the defined UK vantages, respecting site instructions and stop rules.
4. Run fixed discovery and recommendation prompts in fresh sessions; preserve complete output and links.
5. Run safe transaction tasks to the authorised boundary, or record why no valid test was possible.
6. Score indicators from evidence records, not memory or impressions.
7. Independently review applicability, factual extraction, evidence state and score anchor. Adjudicate material disagreements; do not average unsupported interpretations.
8. Reproduce the total and coverage from the frozen dataset. Complete the claims ledger and factual check before publication.

### 8.2 Stop rules

Stop on CAPTCHA escalation, rate-limit warning, suspected harm, login/consent boundary outside the approved scope, unexpected personal-data request, unexpected purchase risk, or any need to bypass a control. Log the stop event and retain permitted evidence. Do not retry in a way that evades the control.

### 8.3 Retailer factual check

Where a retailer result is to be published, share the tested URLs, dates, material evidence and proposed factual description under a fixed response window. Invite evidence corrections, not editorial approval or removal of an unfavourable finding. Evidence introduced after the fieldwork window is labelled as a subsequent change. Log all responses and disclose unresolved material disagreement fairly.

### 8.4 Corrections

Classify changes as typographical, factual non-score, score-affecting or method-affecting. Recalculate every affected comparator where a score or method changes; retain previous versions and publish a dated correction note. Never quietly replace a public result.

## 9. Reporting and claims language

Every public or client report presents:

- AVS total and the six dimension scores;
- publication status, coverage and confidence;
- exact proposition, UK scope, sample, systems and test period;
- access restrictions and unknown/unreachable evidence;
- the three strongest positive findings and three material gaps;
- any difference between declared, observed, sandbox-verified and production-verified capability;
- a clear limitations statement.

Use wording such as:

- “In tests conducted between [dates], [system/version] returned [result] for [n/N] prompts in the UK panel.”
- “The retailer’s robots.txt file restricted [named crawler] on [tested path]; this is a declared crawler rule and does not establish that all AI systems or user-directed actions are blocked.”
- “The product page exposed [field] in the tested source; we did not verify whether every assistant uses that source.”
- “A controlled agent test reached the order-review boundary without placing an order.”

Avoid unqualified claims such as “blocks AI”, “AI cannot buy from them”, “the UK’s most agent-ready retailer”, “agent-proof”, “certified” or “will increase conversion”. State the denominator, panel and period for every quantified claim. Distinguish an announcement from an observed live capability.

## 10. Products using the score

### Agent Visibility Snapshot

An automated preliminary screen. It may show selected Access and Read checks and a limited evidence profile. It must not display the full AVS unless it completes the full six-dimension protocol and publication gates. Label partial results “Snapshot”, not “Score”.

### Agent Visibility Diagnostic

A human-reviewed 21-day assessment using the standard AVS sample plus an expanded product sample, source reconciliation and, where authorised, internal capability evidence. Deliver the comparable outside-in AVS separately from the confidential internal readiness findings and 90-day action plan.

### Agent Visibility Benchmark

A cohort comparison using the same frozen method, category module, system panel, prompt library, fieldwork window and publication gates. Publish only comparable results. Do not rank provisional or ungraded propositions.

### Agent Visibility Monitor

Repeat fixed tests at a declared cadence. Preserve history, display changes against the same panel/method, and flag provider or protocol changes. If the panel or method changes, label a series break or back-test the prior cohort where possible.

## 11. Pilot and validation plan

V1 is ready for a controlled pilot, not for certification or a claim of predictive validity.

### Pilot design

- Five UK retail design partners across materially different categories.
- At least one grocery or convenience proposition, one fashion/footwear proposition, one home/DIY or furniture proposition, one electrical/technology proposition and one health/beauty proposition where safe to assess.
- Run the standard sample twice, with fieldwork windows separated by at least 7 days for the repeatability study, while retaining the normal 24-hour minimum within each scored wave.
- Independently double-score at least 20% of qualitative observations and all disputed material findings.
- Record evaluator time, missingness, score movement, platform instability, cost and evidence availability.
- Interview each partner on interpretability and decision usefulness after the evidence is frozen.

### Pilot exit criteria

Before public league tables or certification are considered, the method owner reviews:

1. whether two reviewers can apply each indicator consistently;
2. whether repeat tests produce stable results or explainable variation;
3. whether the sample is practical to run at the intended cadence;
4. whether the total score changes materially under reasonable indicator or dimension weighting alternatives;
5. whether Access, Read, Trust, Discover, Recommend and Transact each add distinct decision value;
6. whether the existing four-gate index and six-dimension AVS are understood as separate measures;
7. whether the score leads to specific remedial actions without overstating causality.

If weighting sensitivity materially changes a result or ordering, continue to publish the dimension profile and coverage, suspend cross-company ranking and revise the scoring rule under a new method version.

## 12. Version control

- **v1.0** — initial controlled-pilot specification, six dimensions, equal dimension weights, UK retail scope.
- **v1.0.1** — clarification or correction with no scoring effect.
- **v1.1** — new indicator, fieldwork threshold or panel refinement that changes scoring but preserves the construct; back-test where feasible.
- **v2.0** — material change to dimensions, weights, construct, or comparability rules.

Every release records the change, rationale, effective date, affected series and whether historical results were recalculated. Provider token, protocol and product-surface registries receive their own dated versions and are checked before each wave.

## Appendix A. V1 implementation checklist

Before a score can be issued, confirm:

- [ ] UK proposition and seller boundary fixed.
- [ ] Category, product sample and relevance rules fixed.
- [ ] Agent system panel and exact product surfaces frozen.
- [ ] Prompt library and expected-fact sheets frozen.
- [ ] Two access windows and two UK network vantages scheduled.
- [ ] Transaction harness or genuine supported surface declared and authorised.
- [ ] Stop rules and no-purchase boundary communicated to operators.
- [ ] Every test linked to an evidence record and method version.
- [ ] Unknown, blocked, unreachable and not-applicable states kept distinct.
- [ ] Coverage and confidence calculated with the score.
- [ ] Second review and adjudication complete.
- [ ] Comparability and claims gates passed before ranking or publication.

## Appendix B. Primary references

These sources inform the technical interpretation of the method. They do not replace claim-level evidence or provider re-checks before each fieldwork wave.

1. RFC Editor, [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html). The protocol concerns crawler rules and states that those rules are not access authorisation.
2. OpenAI, [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots). Provider documentation distinguishes crawler identities and purposes; re-check the current list and instructions before a wave.
3. Google Search Central, [Product structured data](https://developers.google.com/search/docs/appearance/structured-data/product). Product markup can support Google Search product experiences; its presence alone does not demonstrate that other agents use or interpret the data.
4. The Future Collective, *Agentic Commerce Evidence and Claims Framework: Wave 4 methodology v0.1*, 5 August 2026. Internal companion standard for the four-gate public readiness index, enterprise diagnostic, evidence states, sampling and publication controls.

---

**Method owner:** The Future Collective  
**Next action:** Run the five-retailer pilot, log decisions and deviations, then issue v1.0.1 or v1.1 only if evidence requires it.
