ASRS — Agentic Selling Readiness Score

Rubric v0.4 · how the score works

Reading a score

Five pillars, each scored 0–100 from named checks, then combined by the pillar weights below. Checks that could not be tested shrink their pillar’s denominator — a site is never punished for what couldn’t be observed. Critical failures cap the letter grade regardless of points. Scores are comparable only within a rubric version.

Static checks probe the site directly; BEHAVIORAL checks come from live shopper and trust panels (headless agents working the site under a user directive), including one real zero-value free-tier transaction where the site advertises an allowance.

The rubric, verbatim

# Agentic Selling Readiness Score — rubric v0.4 (2026-07-22)
# The rubric is versioned; every report embeds this version. Scores are only
# comparable within a rubric version (SSL Labs / Euro NCAP pattern).
#
# v0.4 (2026-07-22), two features:
#   (a) Free-tier transaction probe — the first check that actually TRANSACTS
#   instead of doing read-only recon. New bhv_free_tier_transaction (pillar
#   outcome, max 5, behavioral mode only): exercise an ADVERTISED $0 allowance
#   end to end the way a real agent trials a service before spending —
#   discover the free allowance from the target's own docs, make the documented
#   call, receive the HTTP 402 zero-value identity challenge, settle it with a
#   ZERO-VALUE signed authorization from a fresh EPHEMERAL wallet, and verify a
#   real 200 with actual content. Checkpoint partial credit within the check
#   (advertised 1 -> challenge received 1 -> settled $0 2 -> content delivered 1).
#   Safety property: the probe parses the challenge amount and NEVER signs unless
#   it is exactly zero (a nonzero challenge records free-tier-not-zero-cost). No
#   free tier advertised => NA (never punished for its absence). At most ONE
#   attempt per scoring run (it consumes the allowance); --trials never
#   multiplies it. Vendor-neutral: the opt-in header and allowance count are
#   discovered from the agent-surface docs, not hardcoded.
#   (b) Environment-failure attribution + hosted-agent reachability. A shopper
#   run whose OWN hosting stack refused to load the site (hosted URL-safety/
#   reputation layer; self-reported "blocked by browser security policy"
#   language + zero observed checkpoints) is no longer counted as site
#   evidence — it is excluded from outcome and live-trust denominators, and
#   surfaces instead as a new access check, hosted_agent_reachability (max 5):
#   the fraction of verdict-producing runs that could reach the storefront
#   at all.
#
# v0.3 (2026-07-22): rescored for agentic SERVICES, not retail webshops. The
#   organizing question is now "does implementing agent-native rails let an
#   agent do things it otherwise couldn't?": transactability weight 0.20->0.30
#   (access 0.20->0.15, legibility 0.25->0.20); product_schema generalized to
#   offer_catalog (a pricing/catalog manifest JSON or priced OpenAPI counts,
#   not just schema.org JSON-LD); mcp_surface demoted 5->2 (a per-service MCP
#   server duplicates what generic HTTP + a 402 handshake already gives every
#   agent); commerce_protocol_signals folded into x402_probe as partial credit
#   (for services the 402 handshake IS the commerce protocol); self_serve_payg
#   restacked — full marks now require NO-SIGNUP provisioning (signup+API-key
#   self-serve is the middle tier, not the top). (The v0.4 free-tier
#   transaction probe promised here shipped in v0.4 above.)
#
# v0.2 (2026-07-22): x402 probe follows same-domain agent-surface links (the
#   llms.txt breadcrumb to agents.<domain>), reads that surface's discovery
#   docs into the corpus, and elicits the 402 with an empty-body POST where the
#   surface documents x402 — fixes the apex-only false negative (no-x402-surface
#   on a fully x402-payable storefront).
# v0.2 (2026-07-22): trust pillar reframed to DIRECTIVE framing — the panel is
#   asked what it does when the user has already said "go use this site", not
#   whether the site looks scam-like in the abstract. Three-way verdict
#   (proceed / proceed_with_warning / refuse) replaces the binary willing flag;
#   only a confident refusal-despite-directive trips the trust-panel-refusal
#   cap. New trust_live_session check scores the shopper panel's in-task
#   trust_events (warnings surfaced while actually working the site).
# v0.1 (2026-07-22): initial rubric.
version: "0.4"

# Pillar weights for the overall score. If a pillar is entirely
# NA/CANT_TEST (e.g. outcome in static-only mode), it is dropped and the
# remaining weights are renormalized.
pillar_weights:
  access: 0.15          # can an agent get in at all (table stakes)
  legibility: 0.20      # can an agent understand what's for sale
  transactability: 0.30 # can an agent pay + provision WITHOUT a human — the heart of the score
  trust: 0.15           # directed to buy here, does the agent proceed
  outcome: 0.20         # did shopper agents actually get the job done, repeatedly

# Grade caps: critical failures LIMIT the grade regardless of points
# (SSL Labs pattern — averages hide showstoppers). Keyed by finding slug.
caps:
  agent-ua-hard-blocked: 69        # bot wall blocks agent UAs while browsers pass -> max D
  no-https: 59                     # -> F (no HTTPS cannot pass in 2026)
  trust-panel-refusal: 69          # a panel model confidently refuses DESPITE an explicit user directive -> max D
  human-gate-required: 79          # purchase impossible without human-only step -> max C

grade_bands:  # lower bound -> grade
  - [95, "A+"]
  - [90, "A"]
  - [80, "B"]
  - [70, "C"]
  - [60, "D"]
  - [0,  "F"]

checks:
  # ---- access ----
  - id: robots_ai_crawlers
    pillar: access
    max_points: 10
    desc: robots.txt policy for major AI agents (GPTBot, ClaudeBot, Claude-User,
      OAI-SearchBot, PerplexityBot, Google-Extended). Unspecified = allowed.
  - id: agent_ua_reachability
    pillar: access
    max_points: 10
    desc: Fetch key pages with agent user-agents vs a browser UA; detect
      block/challenge (403s, Cloudflare challenge markers) applied only to agents.
  - id: hosted_agent_reachability
    pillar: access
    max_points: 5
    desc: BEHAVIORAL — fraction of shopper runs whose hosting stack allowed
      the site to load at all. Hosted agent platforms gate navigation with
      their own URL-safety/reputation layers; a storefront those layers
      refuse is unreachable to that agent population regardless of its rails.

  # ---- legibility ----
  - id: llms_txt
    pillar: legibility
    max_points: 6
    desc: /llms.txt or /llms-full.txt present and non-trivial — the agent's
      front door, and how agent-surface rails get discovered.
  - id: sitemap
    pillar: legibility
    max_points: 2
    desc: sitemap.xml discoverable (directly or via robots.txt).
  - id: offer_catalog
    pillar: legibility
    max_points: 6
    desc: A machine-readable offer catalog by ANY convention — schema.org
      Product/Offer/Service JSON-LD with price, OR a pricing/catalog manifest
      JSON endpoint (services/meters/plans with prices), OR a priced OpenAPI.
      Services sell metered calls, not SKUs; any parseable catalog counts.
  - id: pricing_machine_readable
    pillar: legibility
    max_points: 4
    desc: A pricing/product page reachable and parseable without JS rendering
      (server-rendered price visible in HTML).
  - id: api_docs_surface
    pillar: legibility
    max_points: 4
    desc: Public docs/API reference discoverable (/docs, /api, openapi spec,
      /.well-known/api-catalog).

  # ---- transactability ----
  - id: x402_probe
    pillar: transactability
    max_points: 8
    desc: Agent-native payment. Full — a live HTTP 402 handshake with a
      payment-requirements payload (x402/MPP) on the site or its agent
      surface. Partial — x402 documented but not elicited, a 402 without a
      parseable payload, or an ACP/UCP commerce-protocol surface. This is the
      strongest "an agent can act here autonomously" proof in the rubric.
  - id: mcp_surface
    pillar: transactability
    max_points: 2
    desc: BONUS — MCP server discoverable (/.well-known/mcp.json, /mcp
      handshake). A per-service MCP server mostly duplicates what generic
      HTTP plus a 402 handshake already gives an agent, so this is worth
      little; its absence is not a defect.
  - id: self_serve_payg
    pillar: transactability
    max_points: 6
    desc: Provisioning tiers. Full — NO-SIGNUP provisioning (an agent can pay
      and call with no account creation; a live 402 handshake proves this).
      Partial — self-serve signup + API key. Fail — sales-gated only. The top
      tier is what lets an agent transact where it otherwise couldn't.

  # ---- trust (static half; behavioral half below) ----
  - id: https_hsts
    pillar: trust
    max_points: 5
    desc: Valid HTTPS; HSTS header present.
  - id: org_identity
    pillar: trust
    max_points: 4
    desc: Verifiable org identity — about/contact/legal pages with a real
      company name and reachable contact.
  - id: policies_present
    pillar: trust
    max_points: 3
    desc: Terms of service, privacy policy, refund/cancellation policy exist.
  - id: reputation_signals
    pillar: trust
    max_points: 3
    desc: Reviews/AggregateRating markup or credible third-party presence.
  - id: trust_panel_willingness
    pillar: trust
    max_points: 8
    desc: BEHAVIORAL — the model panel (Claude, Codex) is told the user has
      DIRECTED it to buy here and asked what it actually does next. Weighted
      per verdict (proceed=1, proceed_with_warning=0.5, refuse=0). Only a
      confident refusal-despite-directive triggers the trust-panel-refusal cap;
      warnings deduct points but never cap.
  - id: trust_live_session
    pillar: trust
    max_points: 4
    desc: BEHAVIORAL — trust as it manifests during the live shopper session
      under a user directive. A run is clean when it reports no trust_events
      (hesitation / would-warn-user moments while working the site); runs with
      events earn half weight. Live evidence (working docs, rails that respond)
      can resolve concerns a static excerpt judge cannot.

  # ---- outcome (behavioral only; NA in static mode) ----
  - id: bhv_found_product
    pillar: outcome
    max_points: 4
    desc: Shopper runs located the product/service matching the task.
  - id: bhv_understood_pricing
    pillar: outcome
    max_points: 3
    desc: Shopper runs extracted concrete price/terms.
  - id: bhv_purchase_path
    pillar: outcome
    max_points: 4
    desc: Shopper runs identified a concrete purchase/provisioning path.
  - id: bhv_machine_payable
    pillar: outcome
    max_points: 5
    desc: That path is machine-payable (API + programmatic payment), not
      browser-checkout-only.
  - id: bhv_no_human_gate
    pillar: outcome
    max_points: 4
    desc: No human-only gate (CAPTCHA, KYC/identity verification, email
      confirmation loop, sales call) blocks the path. The Exa lesson —
      business-rule gates are readiness factors.
  - id: bhv_free_tier_transaction
    pillar: outcome
    max_points: 5
    desc: >-
      BEHAVIORAL — live end-to-end exercise of an ADVERTISED free tier, the one
      check that actually transacts. Discover the free allowance from the
      target's own docs, make the documented call, receive the HTTP 402
      zero-value identity challenge, settle it with a ZERO-VALUE signed
      authorization from a fresh ephemeral wallet, and verify a real 200 with
      content. Checkpoint partial credit — advertised (1), challenge received
      (1), settled $0 (2), real content delivered (1). NA when no free tier is
      advertised (never a defect). It never signs a nonzero authorization (a
      nonzero challenge is free-tier-not-zero-cost). At most ONE attempt per
      scoring run (it consumes the allowance).

← back to the scorecard