Can the named systems reach the intended public information or permitted action route?
Tests named paths and systems. A crawler rule is not proof that all AI systems are blocked.
Controlled-pilot specification · Version 1.0
The Agent Visibility Score (AVS) tests whether a defined UK retail proposition can be reached, understood, verified, found, recommended and progressed towards a safe transaction by a stated panel of AI systems and agent tests.
The method has not yet been externally validated. It makes no claim to predict sales, prove universal agent access or certify that a retailer is ready for every AI system.
Read the full V1.0 pilot specification (Markdown)
The unit is a specific commercial proposition, not a corporate group in the abstract. Before testing, the seller, domain, category, UK market, product sample, systems and observation period are fixed.
Tests named paths and systems. A crawler rule is not proof that all AI systems are blocked.
Checks sampled fields and visible-to-machine consistency, not the mere presence of markup.
Assesses offer truth, policy clarity and source traceability, not brand reputation.
Records dated outcomes for a fixed query set, not universal visibility.
Tests observed recommendations against shopper constraints and checked facts.
Stops before payment or order placement. No purchase is made.
| Dimension | Indicators |
|---|---|
| Access | A1 Crawler policy compatibility; A2 Public endpoint reachability; A3 Content availability through the tested route; A4 Access stability. |
| Read | R1 Entity and variant identity; R2 Decision-relevant field coverage; R3 Machine extraction; R4 Visible-to-machine parity. |
| Trust | T1 Seller and source authority; T2 Offer truth; T3 Policy clarity; T4 Claim and evidence traceability. |
| Discover | D1 Correct entity retrieval; D2 Authoritative source retrieval; D3 Offer and market accuracy. |
| Recommend | M1 Recommendation inclusion rate; M2 Recommendation position; M3 Constraint and fact fit. |
| Transact | X1 Correct selection; X2 Basket and total integrity; X3 Safe progression and hand-off; X4 Recovery and recourse visibility. |
X4 is a pre-purchase visibility check. It does not prove that a completed order can be cancelled, returned or resolved end to end.
AVS = (Access + Read + Trust + Discover + Recommend + Transact) / 6Indicators are averaged within each dimension first. Equal dimension weights are a transparent pilot convention, not a claim of equal commercial impact.
| Sample element | V1.0 standard |
|---|---|
| Retail pages | Three product detail pages and five supporting pages: UK homepage, category, delivery, returns and seller/contact or equivalent. |
| AI system panel | Four declared consumer-facing system slots. The exact product surface, model/version, locale and settings are recorded. Comparable scoring requires at least three available systems. |
| Prompts and runs | Twelve fixed prompts: six discovery and six recommendation missions. Three independent runs per prompt per system, giving 144 response observations for four systems. |
| Access checks | The eight frozen URLs are checked in two fieldwork windows, at least 24 hours apart, from two UK network vantages. |
| Transaction checks | Three safe pre-purchase tasks, attempted twice on the same authorised test surface. The task stops before payment or order placement. |
The initial system panel candidates are ChatGPT with web search, Google Gemini or the relevant Google AI shopping/search surface, Microsoft Copilot and Perplexity. The panel is frozen for each wave; provider changes are documented.
For rate-based indicators, the score is 100 multiplied by successful applicable observations divided by valid applicable observations. Qualitative indicators use anchored scores of 0, 25, 50, 75 or 100, with a written rationale and retained evidence.
Unknown, unreachable or untested evidence is not silently treated as failure. It reduces coverage. A zero is used only when an appropriate test could detect a capability and confirms it absent or materially wrong. All six dimensions need valid observations before a headline AVS can be reported.
| Publication status | Minimum rule | What may be published |
|---|---|---|
| Comparable | At least 80% total evidence coverage, at least 60% in every dimension, three or more systems, the full prompt protocol, transaction tasks attempted and second review complete. | Score and cohort comparison, with sample, coverage and confidence. |
| Provisional | 60-79% total coverage, some valid evidence in every dimension and no material integrity dispute. | Provisional result with gaps explained. No league-table rank or "ready" label. |
| Ungraded | Below 60% coverage, a whole dimension unknown, insufficient panel observations or an unresolved material dispute. | Evidence profile and reason only. No headline score. |
| Confidence | Minimum rule |
|---|---|
| A · High | At least 85% coverage, two or more fieldwork windows, required UK access vantages, complete panel, at least 95% second-review agreement and no material unresolved issue. |
| B · Good | At least 80% coverage, repeat fieldwork and second review complete, with bounded instability disclosed. |
| C · Limited | 60-79% coverage or one important source remains unstable. Provisional presentation only. |
| U · Ungraded | Below 60% coverage, a whole dimension is unknown or a critical integrity issue remains. |
A high score can have low confidence; a low score can have high confidence. V1 reports coverage and confidence alongside the dimension profile. It does not use performance bands, a pass mark, readiness badge, certification or public seal.
| State | Meaning | Scoring treatment |
|---|---|---|
| Confirmed present | An appropriate completed test or reliable primary evidence supports the fact or capability. | Score the observed result. |
| Confirmed absent | A capable test found the item absent or consistently failing. | Score zero only for the relevant indicator. |
| Blocked | A policy, challenge or control prevented the observation. | May affect the tested access indicator; downstream evidence is unknown unless tested through another valid route. |
| Unreachable | A network, server or unresolved technical error prevented a valid observation. | Unknown for scoring; report the friction. |
| Unknown | Evidence cannot establish the position or the route was outside authorisation. | Exclude from score calculation; coverage falls. |
| Not applicable | A written category or business-model rule says the criterion does not apply. | Exclude only with reviewer approval. |
A missing result is not automatically proof that something does not exist. A robots.txt restriction is not a conclusion about every AI agent or user-directed action.
Each report records the exact proposition, sample, systems, observation dates, six dimension scores, coverage, confidence, unknowns and material limitations. Corrections are dated and versioned; a public result is never quietly replaced.
V1.0 is intended for five UK retail design partners across different categories. The pilot repeats the standard assessment, double-scores at least 20% of qualitative observations, and reviews inter-reviewer consistency, repeatability, evaluator time, missingness, cost and decision usefulness.
Before public league tables or certification are considered, the method owner will review reviewer agreement, repeat-test stability, operational practicality, sensitivity to alternative weights, the distinct value of each dimension, whether the four-gate and six-dimension measures are understood separately, and whether findings lead to useful actions without implying causality.
Public league tables and certification are not part of this release. They would require pilot evidence, repeatability analysis, external review and a separately approved method. Scoring or comparability changes are issued as a new, dated version; affected historical comparisons are reviewed.
The instant check reads selected public signals for one domain and one discovered page using three request profiles. It does not run the V1 consumer AI panel, 144 prompt observations, a full product sample, the safe transaction protocol or independent review. Its evidence profile must not be compared with a full six-dimension AVS result.
Run the free SnapshotThe Agentic Shelf Index's four journey gates, Find, Understand, Act and Resolve, are a related but separate research measure. They are not converted into or directly compared with the six-dimension AVS.
| Agentic Shelf Index gate | AVS dimensions that inform it | Boundary |
|---|---|---|
| Find | Access + Discover | AVS tests its named systems and fixed prompt panel. |
| Understand | Read + Trust | AVS tests sampled facts and sources, not every internal data control. |
| Act | Recommend + Transact | AVS tests defined missions and a safe pre-purchase boundary. |
| Resolve | No full AVS equivalent | X4 checks whether recovery and recourse routes are visible; it does not test completed-order resolution end to end. |