← All posts

What x402 Trust Scores Actually Measure

Published 2026-09-04 · the paid-cohort figures below render at build time from the frozen ERC-8004 × x402 ownership join dated 2026-09-01 (snapshot 2026-09-01T11:23:10Z, chain base). The probe-funnel figures are quoted verbatim from the Agents Trust landing page fetched 2026-09-03; their page publishes no snapshot date. All other surface descriptions are frozen against captures dated 2026-09-03 (fuchss, ForgeMesh, x402station) and 2026-08-31 (competitor recensus) — not re-fetched here.

A buyer facing an unfamiliar x402 endpoint has four different questions. The public scoring surfaces answer different ones. Naming which is which is the point of this page.

The one-paragraph answer. A 402 is a payment challenge, never a buyer. There are at least five distinct predicates a scoring surface can measure on an x402 endpoint — advertised, answers a probe, was actually paid, counterparty is a registry-owned identity you can name, and delivered after payment. Every credible surface answers ONE or TWO of them. Naming which layer each surface measures is more useful to a buyer than any single composite score, because the surfaces are complementary layers at different questions, not competing answers to the same question.

The predicate ladder

The layers below are strictly nested from the top: a URL that was paid was necessarily advertised and answers a probe; a URL that answers was necessarily advertised. The reverse implications do not hold, which is why the layers do not substitute for one another.

Layer Predicate on the URL Who measures this at scale
1. Advertised A URL somebody published as an agent endpoint. Registries and catalogs (CDP Bazaar, x402-list.com, Agents Trust, ForgeMesh, others).
2. Answers a probe The URL returns a well-formed payment challenge when asked. This is liveness. Agents Trust (funnel below); ForgeMesh (A–F scan, MPP dual-stack detection).
3. Was actually paid An on-chain settlement was observed against the endpoint's pay_to. x402.fuchss.app (per-endpoint 30d settlements at payTo-wallet level).
4. Counterparty is a registry-owned identity The paid pay_to walks back to an ERC-8004 IdentityRegistry row. The ERC-8004 × x402 ownership join we publish at /x402-paid-but-unrated.
5. Delivered after payment The paid caller actually received the promised resource, without dispute. Nobody, including us — the open frontier.

Layer 5 is called out explicitly because it is the one layer no scoring surface currently addresses; that includes ours. Naming the frontier honestly is more citable than pretending it is already served.

Layer 2 — probe / liveness, measured at scale

Layer 2 answers a question the catalog layer cannot: of the URLs advertised as agent endpoints, which ones actually respond? Answering that at real scale is genuine work. Agents Trust publishes the funnel below on their landing page — we quote it verbatim from a fetch on 2026-09-03 as published on their landing page; no snapshot date is given, so no stronger attribution is available:

Stage URLs Their stated meaning
Advertised to agents 138,204 Catalogued URLs known to their crawl.
Ever answered a probe 64,849 Returned a payment challenge at least once.
Callable now 30,857 Answering a probe on their current live check.

Source: Agents Trust landing page, https://agents-trust.com/, as published on their landing page, fetched 2026-09-03; no snapshot date given. Their predicate is “ask every catalogued URL for a payment challenge” — liveness (does the URL respond to a probe), not settlement (was the URL ever paid, by whom). The landing page also carries an example settlement card built on a fictional URL for illustration; we do not treat that card as an ecosystem statistic and it never appears on this page.

ForgeMesh answers a nearby, narrower question at the endpoint level: if you paste a URL, can an agent actually pay it?. Their scanner returns an A–F grade in about five seconds from an envelope + reachability + paywall-fires probe, and their public census reports: 1 in 4 x402 Bazaar sellers failed a probe-based sale-completion check (ForgeMesh's own August census; predicate = envelope + reachability + paywall-fires). That is a probe-failure census, not a settlement census — the three failure modes are all detectable without any on-chain payment having occurred.

Layer 3 asks whether the endpoint has ever been paid. This is a strictly narrower question than “does the URL respond,” and it is one that x402.fuchss.app answers at scale. Their public product exposes per-endpoint 30-day settlement counts, distinct payer counts, and USDC volume, and their composite trust score combines six named inputs — uptime, envelope compliance, latency, age, on-chain settlement activity, and price stability. They also publish a separate facilitator leaderboard ranking settlement operators by real on-chain USDC volume.

Their scope statement matters and we quote it verbatim, both because it is accurate and because it defines the boundary of this layer:

Settlement figures are wallet-level, not endpoint-specific.” — per-endpoint report note, x402.fuchss.app
We probe and observe; we do not verify delivery-after-payment.” — footer, x402.fuchss.app

Both statements are useful signal to anyone reading a fuchss score: settlement is attributed to the pay_to wallet, which can be shared across many endpoints; and even a successful settlement does not tell you the resource was actually delivered afterwards. Neither caveat undermines the surface — it defines where its measurement stops.

Layer 4 — counterparty is a registry-owned identity

Layer 4 asks a question layer 3 cannot: does the paid pay_to walk back to an on-chain identity you can name? Per-endpoint settlement figures do exist elsewhere (see layer 3). What no other index does is JOIN that paid-endpoint cohort against ERC-8004-owned identity — the on-chain registry that gives the counterparty a name, an agentId, a resolvable metadata URI, and (for the small cohort that has any) a ReputationRegistry row. The whole ownership-join summary lives in one paragraph so the endpoint-level cut can never be read without the publisher-level cut beside it:

The join, both levels, one paragraph. 14,338 earning Base x402 endpoints → 1,291 (9.00%) ERC-8004-owned → 9 rated → 0 commerce-backed; publisher level = 26 distinct pay_tos, top one 67.2% (875 endpoints), 4 with any feedback, 0 commerce-backed. Per-publisher dedup capped at 50 on the base-resolvable count — disclosing this cap on-page is part of the finding, not an addendum.

Sample: 14,338 x402 endpoints on Base with quality.l30DaysUniquePayers >= 1 in the 2026-09-01T11:23:10Z bazaar catalog snapshot. Every share is a lower bound — pay_to may be a treasury wallet distinct from the ERC-8004 registration owner, so the true ownership footprint of the earning cohort is at least this large. Full derivation: /x402-paid-but-unrated; byte-identical projection in the market-map JSON under reputation_join_base.

Nobody else joins the paid-endpoint cohort to on-chain ERC-8004 identity. That is the operation the four surfaces named above do not perform, and it is what this layer answers.

Undetermined at capture: x402station

x402station is included here for completeness. The methodology narrative on their site (x402station.com, fetched 2026-09-03) describes a HYBRID surface built on real transaction volumes, success rates, and unique users. The live site, however, shipped 0 x402 Services Available, 0 Total Transactions, and 0.0% Average Reliability at capture — so the stated methodology could not be inspected because there was nothing operational to inspect. Marked UNDETERMINED. Whether the surface eventually ships data with the promised shape is a question for a later re-fetch, not this piece.

Closed-loop vs registry-derived scoring

A recurring architectural choice worth naming plainly. A score bootstrapped only from payments the scorer itself mediates cold-starts at zero and cannot describe an agent it has never seen. A registry-derived score exists whether or not the agent ever transacted with the scorer — because the score is derived from an on-chain identity the agent already has. The two shapes are not competitors; they answer different questions. Closed-loop scoring is deep on the endpoints inside the loop and silent everywhere else. Registry-derived scoring is broad across the identity graph and shallower per-endpoint.

In the 2026-08-31 competitor recensus (docs/evidence/0369-competitor-recensus-2026-08-31.md), eight distinct entrants in the counterparty-scoring space were verified at source. Every one of the five items in that pass with a live landing page shipped zero disclosed adoption — no user count, no scored-agent count, no call count on any page fetched. “Disclosed” is doing load-bearing work in that sentence: absence of a public metric is not evidence the product is unused, only that adoption has not been declared.

A 402 is a challenge, never a buyer

The hinge of this whole piece. When a URL responds with an HTTP 402 Payment Required, that is a challenge: it advertises the terms of payment. It says nothing about whether anybody has ever met those terms. An endpoint can serve a well-formed 402 for years without ever being paid once. Probe-liveness measurement — layer 2 above — is a genuinely useful, complementary layer of signal; it just answers a different question from the paid layer.

If you are about to pay, run these checks

Everything above is a map of what the public surfaces measure. The buyer-side checklist — the actual sequence of checks against public data before you send money — lives on /blog/how-to-vet-an-x402-endpoint. In short:

  1. Ask whether the endpoint has been paid at all, and by how many distinct payers — the payer-count cohort ladder on /x402-bazaar-market-map answers this from the CDP Bazaar snapshot.
  2. Ask whether the traffic is diffuse or concentrated — a one-payer endpoint is indistinguishable from self-dealing on-chain. The exclusive bucket distribution on the market map shows the shape.
  3. Ask whether the endpoint's pay_to walks back to a registry-owned identity — /x402-paid-but-unrated is the reference join.
  4. Ask whether the same publisher operates ten more endpoints — the publisher-level cut in the join paragraph above sits beside the endpoint-level cut for exactly this reason.
  5. Ask whether any of it carries commerce-backed reputation — the count above is the current answer for the earning Base cohort. See also /blog/probe-vs-paid for why layers 2 and 3 are strictly nested and cannot substitute for one another.

Methodology and scope

External figures. The Agents Trust funnel (advertised → ever answered → callable now) is quoted verbatim from https://agents-trust.com/ as published on their landing page, fetched 2026-09-03 — their page publishes no snapshot date. The ForgeMesh “1 in 4” figure is their own August census, quoted from forgemesh.io/scan as captured 2026-09-03; predicate is envelope + reachability + paywall-fires, not settlement. No fuchss numeric figures are quoted anywhere on this page — only their two verbatim scope statements. The x402station line is a live-fetch capture from x402station.com dated 2026-09-03 (the .io domain does not resolve).

Our figures. The endpoint- and publisher-level ownership-join numbers come from the frozen ERC-8004 × x402 join dated 2026-09-01 (snapshot 2026-09-01T11:23:10Z), predicate smartcontractauditpro/commerce_backed.py. Full derivation and per-publisher figures live at /x402-paid-but-unrated; byte-identical projection in the market-map JSON under reputation_join_base. Per-publisher dedup cap of 50 applies to the base-resolvable count on that join, disclosed in the paragraph above.

Cross-cutting discipline. l30DaysTotalCalls from the CDP Bazaar is a request count, not settled USDC value. The CDP Bazaar catalog is one Base-weighted x402 discovery source; the probe funnel is one operator's crawl over what they catalogue. Neither is an ecosystem census. Every share we quote on the join is a lower bound — pay_to may be a treasury wallet distinct from the on-chain registration owner, and the ERC-8004 spec is still Draft.

Suggested citation: On-Chain Agent Intel. "What x402 Trust Scores Actually Measure (2026-09-04)." https://onchainagentintel.io/blog/what-x402-trust-scores-actually-measure. Published under CC BY 4.0 — attribution back to On-Chain Agent Intel is required.