THE FUTURE COLLECTIVE

Controlled-pilot specification · Version 1.0

Agent Visibility Score methodology

The Agent Visibility Score (AVS) tests whether a defined UK retail proposition can be reached, understood, verified, found, recommended and progressed towards a safe transaction by a stated panel of AI systems and agent tests.

V1.0 is ready for a controlled pilot, not public certification.

The method has not yet been externally validated. It makes no claim to predict sales, prove universal agent access or certify that a retailer is ready for every AI system.

6 dimensionsEqual weight, 0 to 100 each
UK retailInitial market and test scope
Evidence-ledEvery result carries coverage and confidence

What the score measures

The unit is a specific commercial proposition, not a corporate group in the abstract. Before testing, the seller, domain, category, UK market, product sample, systems and observation period are fixed.

01 · Access

Can the named systems reach the intended public information or permitted action route?

Tests named paths and systems. A crawler rule is not proof that all AI systems are blocked.

02 · Read

Can the right product, variant, attributes and offer be interpreted?

Checks sampled fields and visible-to-machine consistency, not the mere presence of markup.

03 · Trust

Can an agent verify who sells the item and where material facts come from?

Assesses offer truth, policy clarity and source traceability, not brand reputation.

04 · Discover

Do named systems find the relevant organisation, product or authoritative source?

Records dated outcomes for a fixed query set, not universal visibility.

05 · Recommend

Does the proposition appear accurately in defined shopping missions?

Tests observed recommendations against shopper constraints and checked facts.

06 · Transact

Can an authorised agent progress a task to a safe review or hand-off?

Stops before payment or order placement. No purchase is made.

View the 22 core indicators
DimensionIndicators
AccessA1 Crawler policy compatibility; A2 Public endpoint reachability; A3 Content availability through the tested route; A4 Access stability.
ReadR1 Entity and variant identity; R2 Decision-relevant field coverage; R3 Machine extraction; R4 Visible-to-machine parity.
TrustT1 Seller and source authority; T2 Offer truth; T3 Policy clarity; T4 Claim and evidence traceability.
DiscoverD1 Correct entity retrieval; D2 Authoritative source retrieval; D3 Offer and market accuracy.
RecommendM1 Recommendation inclusion rate; M2 Recommendation position; M3 Constraint and fact fit.
TransactX1 Correct selection; X2 Basket and total integrity; X3 Safe progression and hand-off; X4 Recovery and recourse visibility.

X4 is a pre-purchase visibility check. It does not prove that a completed order can be cancelled, returned or resolved end to end.

Headline calculationAVS = (Access + Read + Trust + Discover + Recommend + Transact) / 6

Indicators are averaged within each dimension first. Equal dimension weights are a transparent pilot convention, not a claim of equal commercial impact.

How a standard assessment works

  1. Freeze the scopeRecord the exact retailer proposition, seller, UK domain, category, products, systems, prompts and method version before results are visible.
  2. Capture public evidenceTest the frozen product and supporting pages, preserving request details, returned content, timestamps, source facts and any stop events.
  3. Run the fixed system panelUse the same discovery and recommendation prompts, runs and UK settings for every proposition in a comparable cohort.
  4. Score and reviewScore from evidence records, calculate coverage and confidence, complete a second review, and adjudicate material disagreement before publication.
Sample elementV1.0 standard
Retail pagesThree product detail pages and five supporting pages: UK homepage, category, delivery, returns and seller/contact or equivalent.
AI system panelFour declared consumer-facing system slots. The exact product surface, model/version, locale and settings are recorded. Comparable scoring requires at least three available systems.
Prompts and runsTwelve fixed prompts: six discovery and six recommendation missions. Three independent runs per prompt per system, giving 144 response observations for four systems.
Access checksThe eight frozen URLs are checked in two fieldwork windows, at least 24 hours apart, from two UK network vantages.
Transaction checksThree safe pre-purchase tasks, attempted twice on the same authorised test surface. The task stops before payment or order placement.

The initial system panel candidates are ChatGPT with web search, Google Gemini or the relevant Google AI shopping/search surface, Microsoft Copilot and Perplexity. The panel is frozen for each wave; provider changes are documented.

Scoring, missing evidence and comparability

For rate-based indicators, the score is 100 multiplied by successful applicable observations divided by valid applicable observations. Qualitative indicators use anchored scores of 0, 25, 50, 75 or 100, with a written rationale and retained evidence.

Unknown, unreachable or untested evidence is not silently treated as failure. It reduces coverage. A zero is used only when an appropriate test could detect a capability and confirms it absent or materially wrong. All six dimensions need valid observations before a headline AVS can be reported.

Publication statusMinimum ruleWhat may be published
ComparableAt least 80% total evidence coverage, at least 60% in every dimension, three or more systems, the full prompt protocol, transaction tasks attempted and second review complete.Score and cohort comparison, with sample, coverage and confidence.
Provisional60-79% total coverage, some valid evidence in every dimension and no material integrity dispute.Provisional result with gaps explained. No league-table rank or "ready" label.
UngradedBelow 60% coverage, a whole dimension unknown, insufficient panel observations or an unresolved material dispute.Evidence profile and reason only. No headline score.

Confidence grades

ConfidenceMinimum rule
A · HighAt least 85% coverage, two or more fieldwork windows, required UK access vantages, complete panel, at least 95% second-review agreement and no material unresolved issue.
B · GoodAt least 80% coverage, repeat fieldwork and second review complete, with bounded instability disclosed.
C · Limited60-79% coverage or one important source remains unstable. Provisional presentation only.
U · UngradedBelow 60% coverage, a whole dimension is unknown or a critical integrity issue remains.
Confidence is separate from performance.

A high score can have low confidence; a low score can have high confidence. V1 reports coverage and confidence alongside the dimension profile. It does not use performance bands, a pass mark, readiness badge, certification or public seal.

Evidence states

StateMeaningScoring treatment
Confirmed presentAn appropriate completed test or reliable primary evidence supports the fact or capability.Score the observed result.
Confirmed absentA capable test found the item absent or consistently failing.Score zero only for the relevant indicator.
BlockedA policy, challenge or control prevented the observation.May affect the tested access indicator; downstream evidence is unknown unless tested through another valid route.
UnreachableA network, server or unresolved technical error prevented a valid observation.Unknown for scoring; report the friction.
UnknownEvidence cannot establish the position or the route was outside authorisation.Exclude from score calculation; coverage falls.
Not applicableA written category or business-model rule says the criterion does not apply.Exclude only with reviewer approval.

A missing result is not automatically proof that something does not exist. A robots.txt restriction is not a conclusion about every AI agent or user-directed action.

Safety, claims and corrections

No purchasePublic transaction checks stop before payment or order placement. Private routes require explicit authorisation and an agreed test environment.
No bypassing controlsTesting stops on CAPTCHA escalation, rate limits, unexpected personal-data requests or purchase risk.
No universal claimsResults describe named systems, prompts, pages, dates and locations. They do not prove that a retailer "blocks AI".

Each report records the exact proposition, sample, systems, observation dates, six dimension scores, coverage, confidence, unknowns and material limitations. Corrections are dated and versioned; a public result is never quietly replaced.

Controlled pilot and versioning

V1.0 is intended for five UK retail design partners across different categories. The pilot repeats the standard assessment, double-scores at least 20% of qualitative observations, and reviews inter-reviewer consistency, repeatability, evaluator time, missingness, cost and decision usefulness.

Before public league tables or certification are considered, the method owner will review reviewer agreement, repeat-test stability, operational practicality, sensitivity to alternative weights, the distinct value of each dimension, whether the four-gate and six-dimension measures are understood separately, and whether findings lead to useful actions without implying causality.

Public league tables and certification are not part of this release. They would require pilot evidence, repeatability analysis, external review and a separately approved method. Scoring or comparability changes are issued as a new, dated version; affected historical comparisons are reviewed.

How this relates to the site tools

The free tool is a Snapshot, not a full AVS assessment

The instant check reads selected public signals for one domain and one discovered page using three request profiles. It does not run the V1 consumer AI panel, 144 prompt observations, a full product sample, the safe transaction protocol or independent review. Its evidence profile must not be compared with a full six-dimension AVS result.

Run the free Snapshot

The Agentic Shelf Index's four journey gates, Find, Understand, Act and Resolve, are a related but separate research measure. They are not converted into or directly compared with the six-dimension AVS.

Agentic Shelf Index gateAVS dimensions that inform itBoundary
FindAccess + DiscoverAVS tests its named systems and fixed prompt panel.
UnderstandRead + TrustAVS tests sampled facts and sources, not every internal data control.
ActRecommend + TransactAVS tests defined missions and a safe pre-purchase boundary.
ResolveNo full AVS equivalentX4 checks whether recovery and recourse routes are visible; it does not test completed-order resolution end to end.