Rubric v0.4 · how the score works
Five pillars, each scored 0–100 from named checks, then combined by the pillar weights below. Checks that could not be tested shrink their pillar’s denominator — a site is never punished for what couldn’t be observed. Critical failures cap the letter grade regardless of points. Scores are comparable only within a rubric version.
Static checks probe the site directly; BEHAVIORAL checks come from live shopper and trust panels (headless agents working the site under a user directive), including one real zero-value free-tier transaction where the site advertises an allowance.
# Agentic Selling Readiness Score — rubric v0.4 (2026-07-22)
# The rubric is versioned; every report embeds this version. Scores are only
# comparable within a rubric version (SSL Labs / Euro NCAP pattern).
#
# v0.4 (2026-07-22), two features:
# (a) Free-tier transaction probe — the first check that actually TRANSACTS
# instead of doing read-only recon. New bhv_free_tier_transaction (pillar
# outcome, max 5, behavioral mode only): exercise an ADVERTISED $0 allowance
# end to end the way a real agent trials a service before spending —
# discover the free allowance from the target's own docs, make the documented
# call, receive the HTTP 402 zero-value identity challenge, settle it with a
# ZERO-VALUE signed authorization from a fresh EPHEMERAL wallet, and verify a
# real 200 with actual content. Checkpoint partial credit within the check
# (advertised 1 -> challenge received 1 -> settled $0 2 -> content delivered 1).
# Safety property: the probe parses the challenge amount and NEVER signs unless
# it is exactly zero (a nonzero challenge records free-tier-not-zero-cost). No
# free tier advertised => NA (never punished for its absence). At most ONE
# attempt per scoring run (it consumes the allowance); --trials never
# multiplies it. Vendor-neutral: the opt-in header and allowance count are
# discovered from the agent-surface docs, not hardcoded.
# (b) Environment-failure attribution + hosted-agent reachability. A shopper
# run whose OWN hosting stack refused to load the site (hosted URL-safety/
# reputation layer; self-reported "blocked by browser security policy"
# language + zero observed checkpoints) is no longer counted as site
# evidence — it is excluded from outcome and live-trust denominators, and
# surfaces instead as a new access check, hosted_agent_reachability (max 5):
# the fraction of verdict-producing runs that could reach the storefront
# at all.
#
# v0.3 (2026-07-22): rescored for agentic SERVICES, not retail webshops. The
# organizing question is now "does implementing agent-native rails let an
# agent do things it otherwise couldn't?": transactability weight 0.20->0.30
# (access 0.20->0.15, legibility 0.25->0.20); product_schema generalized to
# offer_catalog (a pricing/catalog manifest JSON or priced OpenAPI counts,
# not just schema.org JSON-LD); mcp_surface demoted 5->2 (a per-service MCP
# server duplicates what generic HTTP + a 402 handshake already gives every
# agent); commerce_protocol_signals folded into x402_probe as partial credit
# (for services the 402 handshake IS the commerce protocol); self_serve_payg
# restacked — full marks now require NO-SIGNUP provisioning (signup+API-key
# self-serve is the middle tier, not the top). (The v0.4 free-tier
# transaction probe promised here shipped in v0.4 above.)
#
# v0.2 (2026-07-22): x402 probe follows same-domain agent-surface links (the
# llms.txt breadcrumb to agents.<domain>), reads that surface's discovery
# docs into the corpus, and elicits the 402 with an empty-body POST where the
# surface documents x402 — fixes the apex-only false negative (no-x402-surface
# on a fully x402-payable storefront).
# v0.2 (2026-07-22): trust pillar reframed to DIRECTIVE framing — the panel is
# asked what it does when the user has already said "go use this site", not
# whether the site looks scam-like in the abstract. Three-way verdict
# (proceed / proceed_with_warning / refuse) replaces the binary willing flag;
# only a confident refusal-despite-directive trips the trust-panel-refusal
# cap. New trust_live_session check scores the shopper panel's in-task
# trust_events (warnings surfaced while actually working the site).
# v0.1 (2026-07-22): initial rubric.
version: "0.4"
# Pillar weights for the overall score. If a pillar is entirely
# NA/CANT_TEST (e.g. outcome in static-only mode), it is dropped and the
# remaining weights are renormalized.
pillar_weights:
access: 0.15 # can an agent get in at all (table stakes)
legibility: 0.20 # can an agent understand what's for sale
transactability: 0.30 # can an agent pay + provision WITHOUT a human — the heart of the score
trust: 0.15 # directed to buy here, does the agent proceed
outcome: 0.20 # did shopper agents actually get the job done, repeatedly
# Grade caps: critical failures LIMIT the grade regardless of points
# (SSL Labs pattern — averages hide showstoppers). Keyed by finding slug.
caps:
agent-ua-hard-blocked: 69 # bot wall blocks agent UAs while browsers pass -> max D
no-https: 59 # -> F (no HTTPS cannot pass in 2026)
trust-panel-refusal: 69 # a panel model confidently refuses DESPITE an explicit user directive -> max D
human-gate-required: 79 # purchase impossible without human-only step -> max C
grade_bands: # lower bound -> grade
- [95, "A+"]
- [90, "A"]
- [80, "B"]
- [70, "C"]
- [60, "D"]
- [0, "F"]
checks:
# ---- access ----
- id: robots_ai_crawlers
pillar: access
max_points: 10
desc: robots.txt policy for major AI agents (GPTBot, ClaudeBot, Claude-User,
OAI-SearchBot, PerplexityBot, Google-Extended). Unspecified = allowed.
- id: agent_ua_reachability
pillar: access
max_points: 10
desc: Fetch key pages with agent user-agents vs a browser UA; detect
block/challenge (403s, Cloudflare challenge markers) applied only to agents.
- id: hosted_agent_reachability
pillar: access
max_points: 5
desc: BEHAVIORAL — fraction of shopper runs whose hosting stack allowed
the site to load at all. Hosted agent platforms gate navigation with
their own URL-safety/reputation layers; a storefront those layers
refuse is unreachable to that agent population regardless of its rails.
# ---- legibility ----
- id: llms_txt
pillar: legibility
max_points: 6
desc: /llms.txt or /llms-full.txt present and non-trivial — the agent's
front door, and how agent-surface rails get discovered.
- id: sitemap
pillar: legibility
max_points: 2
desc: sitemap.xml discoverable (directly or via robots.txt).
- id: offer_catalog
pillar: legibility
max_points: 6
desc: A machine-readable offer catalog by ANY convention — schema.org
Product/Offer/Service JSON-LD with price, OR a pricing/catalog manifest
JSON endpoint (services/meters/plans with prices), OR a priced OpenAPI.
Services sell metered calls, not SKUs; any parseable catalog counts.
- id: pricing_machine_readable
pillar: legibility
max_points: 4
desc: A pricing/product page reachable and parseable without JS rendering
(server-rendered price visible in HTML).
- id: api_docs_surface
pillar: legibility
max_points: 4
desc: Public docs/API reference discoverable (/docs, /api, openapi spec,
/.well-known/api-catalog).
# ---- transactability ----
- id: x402_probe
pillar: transactability
max_points: 8
desc: Agent-native payment. Full — a live HTTP 402 handshake with a
payment-requirements payload (x402/MPP) on the site or its agent
surface. Partial — x402 documented but not elicited, a 402 without a
parseable payload, or an ACP/UCP commerce-protocol surface. This is the
strongest "an agent can act here autonomously" proof in the rubric.
- id: mcp_surface
pillar: transactability
max_points: 2
desc: BONUS — MCP server discoverable (/.well-known/mcp.json, /mcp
handshake). A per-service MCP server mostly duplicates what generic
HTTP plus a 402 handshake already gives an agent, so this is worth
little; its absence is not a defect.
- id: self_serve_payg
pillar: transactability
max_points: 6
desc: Provisioning tiers. Full — NO-SIGNUP provisioning (an agent can pay
and call with no account creation; a live 402 handshake proves this).
Partial — self-serve signup + API key. Fail — sales-gated only. The top
tier is what lets an agent transact where it otherwise couldn't.
# ---- trust (static half; behavioral half below) ----
- id: https_hsts
pillar: trust
max_points: 5
desc: Valid HTTPS; HSTS header present.
- id: org_identity
pillar: trust
max_points: 4
desc: Verifiable org identity — about/contact/legal pages with a real
company name and reachable contact.
- id: policies_present
pillar: trust
max_points: 3
desc: Terms of service, privacy policy, refund/cancellation policy exist.
- id: reputation_signals
pillar: trust
max_points: 3
desc: Reviews/AggregateRating markup or credible third-party presence.
- id: trust_panel_willingness
pillar: trust
max_points: 8
desc: BEHAVIORAL — the model panel (Claude, Codex) is told the user has
DIRECTED it to buy here and asked what it actually does next. Weighted
per verdict (proceed=1, proceed_with_warning=0.5, refuse=0). Only a
confident refusal-despite-directive triggers the trust-panel-refusal cap;
warnings deduct points but never cap.
- id: trust_live_session
pillar: trust
max_points: 4
desc: BEHAVIORAL — trust as it manifests during the live shopper session
under a user directive. A run is clean when it reports no trust_events
(hesitation / would-warn-user moments while working the site); runs with
events earn half weight. Live evidence (working docs, rails that respond)
can resolve concerns a static excerpt judge cannot.
# ---- outcome (behavioral only; NA in static mode) ----
- id: bhv_found_product
pillar: outcome
max_points: 4
desc: Shopper runs located the product/service matching the task.
- id: bhv_understood_pricing
pillar: outcome
max_points: 3
desc: Shopper runs extracted concrete price/terms.
- id: bhv_purchase_path
pillar: outcome
max_points: 4
desc: Shopper runs identified a concrete purchase/provisioning path.
- id: bhv_machine_payable
pillar: outcome
max_points: 5
desc: That path is machine-payable (API + programmatic payment), not
browser-checkout-only.
- id: bhv_no_human_gate
pillar: outcome
max_points: 4
desc: No human-only gate (CAPTCHA, KYC/identity verification, email
confirmation loop, sales call) blocks the path. The Exa lesson —
business-rule gates are readiness factors.
- id: bhv_free_tier_transaction
pillar: outcome
max_points: 5
desc: >-
BEHAVIORAL — live end-to-end exercise of an ADVERTISED free tier, the one
check that actually transacts. Discover the free allowance from the
target's own docs, make the documented call, receive the HTTP 402
zero-value identity challenge, settle it with a ZERO-VALUE signed
authorization from a fresh ephemeral wallet, and verify a real 200 with
content. Checkpoint partial credit — advertised (1), challenge received
(1), settled $0 (2), real content delivered (1). NA when no free tier is
advertised (never a defect). It never signs a nonzero authorization (a
nonzero challenge is free-tier-not-zero-cost). At most ONE attempt per
scoring run (it consumes the allowance).