The Truman Show Economy: What AI Agents See When They Try to Buy

By philpher0x

85% of x402 payments aren’t what they look like. AI agents see payments, reviews, rankings, and providers, but still struggle to know what’s real. ERC-8004 adds public reputation — but that reputation is easy to manipulate. And now we have a new candidate, 402Pilot: another early attempt to help AI agents escape this “Truman Show” when deciding what — and who — to pay.

Tags: x402, Микроплатежи, Микроплатежи, 402Pilot, ERC-8004, Микроплатежи, Микроплатежи

85% of x402 payments aren’t what they look like.

AI agents see payments, reviews, rankings, and providers, but still struggle to know what’s real. ERC-8004 adds public reputation — but that reputation is easy to manipulate. And now we have a new candidate, 402Pilot: another early attempt to help AI agents escape this “Truman Show” when deciding what — and who — to pay.

The agentic economy and machine payments have problems few people talk about. When I went through the resources of the Coinbase facilitator, the results surprised me. Plenty of sellers, few buyers and few sales; most of the services are just a long tail with zero or one sale — enough to land in the Bazaar catalog and nothing more, with no real customers at all — plus faked metrics and catalog spam coming from a single domain.

The situation looks like the internet of the '90s or the early 2000s, when the rules and the censorship were only being set up.

In practice, AI agents have nothing to lean on when they assess sellers and their services. AI agents are very easy to fool. A well-crafted description for better semantic search, inflated metrics and fake reviews for reputation — and the agent decides the service is worth its attention.

Reputation and decentralized tooling like the ERC-8004 standard don't help yet either, for exactly the same reason: reputation gets gamed.

Right now I see this as the main problem and the main obstacle to the healthy development of the agentic economy.

An agent can solve a task, and as long as it does that for free, it can pull from many sources and data providers: the client only pays for the LLM, and a bad result can be redone. But once the data becomes paid and the agent can't pick the best and most trustworthy source, letting it make payments at all is simply pointless.

The market as AI sees it

A team from City University of Hong Kong in paper "How Agentic Is Agentic Commerce?" produced the first population-scale measurement of x402 on Base over a 280-day window: 136,708,672 payments worth $44,121,383.81. It sounds like a functioning economy.

They built a payment graph and showed what the payments actually are:

Payment class Share What it means
Fictitious 21.20% The funds never leave the sender — paying yourself
Internal (inside a linked cluster) 63.78% The operator settles between its own addresses
Unattributed 15.02% No proof of independence, and no proof against — these look the most like the real ones

84.98% of settlements are operators settling with themselves.

A huge fictitious layer in which, on top of everything else, the customer bases barely overlap either. Which means the customers belong to the same address clusters as well.

The study's data on the Base facilitator matches mine almost exactly: 25,163 resources collapse into 811 recipient addresses. Of those, 624 ever received a payment, and 249 earned at least $10 over their entire lifetime.

Hence the first conclusion: an AI agent walks not into a market, but into the stage set of one.

The storefront, the popularity counters and the "payment volume" are produced by the same side that does the selling. And almost for free: reproducing the entire array of 136 million payments costs around $355,583 in gas.

As you can see, the data differs sharply from what gets posted on Twitter. The volume and transaction counts may well match — but none of it means anything if sellers are paying themselves.

Reputation: the layer that was supposed to close this

The logical answer is reputation. Let buyers leave reviews, and let the next buyer read them.

ERC-8004 is the first attempt to do that without intermediaries: three on-chain registries — Identity, Reputation and Validation. Ethereum mainnet in January 2026, then BSC and Base. By May: more than 170,000 registered agents and more than 150,000 feedback records.

In a paper "Can Trustless Agents Be Trusted" a group of researchers collected every registry event across the three networks and checked whether anything can be decided on top of it.

The answer: no. Or at least — not yet.

Here is what that study found:

What they checked What they found
Who writes the reviews On BSC, 76 addresses wrote 29,444 reviews — about 387 reviews per reviewer (40 on Base, 5 on Ethereum)
Whether there is any interaction with the resource behind a review No proof of payment or task behind ~98% of reviews
Whether the reviewers ever paid for anything Only 6.2% of reviewers have any payment history at all
What it costs to convince an agent that a resource can be trusted (gas for the transactions) $0.055 (Ethereum) / $0.0027 (Base)
Share of Sybil reviews 41.4% (Ethereum) / 96.3% (BSC) / 92.6% (Base)
Left without a single valid review after cleanup 15.8% / 77.9% / 86.8%

The Validation Registry — the registry of stake-backed proofs meant for expensive deals — was never deployed on mainnet in any of the three networks.

And the final detail that finishes off the idea of decentralized reputation: on Ethereum, the most heavily gamed agents are exactly the ones with a valid registration file and a declared service — 61.1% Sybil reviews. People game whoever has something worth gaming.

To be fair: the authors criticize the current deployment, not the idea. ERC-8004 does exactly what it promises — public identities and public reviews — while deliberately leaving semantics, proofs and Sybil resistance to off-chain systems for the sake of easy adoption. The problem is that the buyer needs a solution today, and those systems barely exist.

As for discovery: metadata manipulation and Sybil flooding in the catalog are not a hypothesis. Five Attacks on x402 describes a dedicated attack on service selection in which a single server intercepted 71.8% of agent traffic. You can see roughly the same thing in the Coinbase facilitator's resources, where one server added almost half of all listings (as of the time of the research).

Why any external check breaks down

To produce a positive signal about itself, a seller needs N cheap actions: spin up wallets, buy from itself, rate itself. On Base that's a fraction of a cent per action.

To refute that signal, you need a real buyer who paid the seller with their own money and verified the result. The cost of gaming is linear and small; the cost of the truth is the price of a real purchase plus the price of verification.

On a thin market that asymmetry breaks everything. When the category leader has 200 organic payers, five hundred fake wallets can simply flip the ranking.

That is exactly why in Coinbase's Bazaar catalog 20 domains hold 58% of the entire "unique payers" metric: not because anyone is particularly inventive, but because the cost of entry into gaming is lower than the cost of entry into honest demand.

Next comes Goodhart's law. Any public criterion agents use to pick a seller instantly becomes the sellers' optimization target. Bazaar's "unique payers over 30 days" metric was sensible right up until the day it started driving rankings. The average score in ERC-8004 was sensible until exactly the same moment.

Reputation answers the question "is this service good in general." What the buyer needs is an answer to "is it good for my task, in my format, at my volumes, today." A provider that's excellent on short factual queries can fall apart on multi-step ones. A provider that was the best yesterday has degraded under load today. This is contextual, shifting knowledge — an averaged number doesn't reflect it.

Which brings me to a conclusion I'm not entirely comfortable with: the most reliable thing for a buyer is to run its own resource validator. Uncomfortable because it complicates the architecture of agents and the way they interact with the world.

Unfortunately, a general rating or an industry-wide score can be a poor fit for a specific agent's goals.

And privacy here isn't paranoia, it's a security property. A public criterion can be studied and optimized against without improving anything in substance.

A criterion the seller can't see can't be gamed — the only way to score on it is to actually do the work well.

That's the one thing that will genuinely work in the buyer's favor.

402Pilot: the first framing of the problem

And this is where one more recent paper comes in: in early August 2026 a team of researchers published new paper 402Pilot: An x402 Decision Layer for Autonomous Agent Micropayments. 402Pilot — not a protocol and not a wallet, but an attempt to describe a decision layer on the buyer's side.

x402 answers the question "how do I move the money." 402Pilot answers "to whom, and how much."

The authors define the problem by five conditions at once:

  1. the wallet balance is finite, and the money really does run out;
  2. the payment is irreversible and happens before the result arrives;
  3. feedback comes only after the result;
  4. prices and server responses change without warning;
  5. the price is learned from the receipt, not from the price list.

It's this set of five conditions that separates the problem from model routing. FrugalGPT, RouteLLM, OpenRouter and LiteLLM all choose within a known menu with a price list you already have subscription access to. None of them learns from the amount actually charged for a single call.

How 402Pilot is put together:

402Pilot is a decision layer between the AI agent and x402 payments that decides which of the available paid providers a given request should go to. It takes into account the task context, service prices, the remaining budget and the burn rate, after which the PA-DCT policy picks a provider; x402 executes the payment and the request, and 402Pilot scores the result it received along with its cost and uses that feedback for the next decisions. This way the system gradually learns when it's worth paying more for quality and when to take the cheap service, while adapting to changes in providers' prices and quality.

For the provider-selection policy the authors use their own term: PA-DCT (Payment-Aware Discounted Contextual Thompson Sampling)

The PA-DCT policy in plain language:

Thompson sampling. The agent doesn't store average scores for providers. It stores "what I believe and how unsure I am," and before every purchase it draws a random guess out of that belief and picks the best of the draws. A provider it knows little about has a wide spread — and wins the draw every now and then. That's how market exploration happens by itself, without a separate budget for experiments.

Wallet pressure. The system compares the actual burn rate against a linear plan. On plan — price and quality carry equal weight. Burning faster — price starts to outweigh, and the expensive tiers drop out of consideration. Saving — premium opens up.

Forgetting. All accumulated observations are discounted every round, including observations about providers the agent isn't currently buying from. The side effect turned out to be the main function: a provider the agent turned away from is gradually "forgotten" back to prior uncertainty, starts winning draws again — and the agent notices that a service that went down has been fixed. Without that mechanism, one failure would be a life sentence.

The 402Pilot-Bench benchmark deserves a separate mention: 823 tasks from HumanEval, HotpotQA, TriviaQA and OpenAssistant, five providers, 20,575 pre-scored responses, three market regimes, 30 paired seeds, 10,000 rounds, a $50 wallet.

Provider Price What it does
P-cheap $0.0005 Cheap model, mediocre quality
P-mid $0.002 Honest mid tier
P-premium $0.01 Expensive model, best quality
P-adv $0.002 Fluently produces a wrong answer 60% of the time
P-flaky $0.002 Bills a timeout 40% of the time

Three providers with the same price and the same base model are indistinguishable on the price list. The difference only shows up after the payment. The model bakes in exactly the two defects that appear in no price list and that no reputation system catches: fake quality and paid non-delivery.

P-adv is essentially an economic model of the service-selection attack from the previous section.

Unfortunately, in a calm market the learning policy loses to the dumb "always take the mid tier" strategy. Learning pays off only where the market moves: a provider breaks, a price drops, quality slips.

What 402Pilot doesn't solve

The authors list the limitations themselves:

  • There's no live traffic. All the numbers come from replaying a cache of pre-generated responses.
  • Discovery is assumed solved. The candidate list and the quotes arrive from outside. Where they come from and whether the metadata can be trusted is out of scope.
  • Sellers aren't strategic. Prices are posted, nobody adapts to the buyer, there's no bargaining and no auctions.
  • No retries. One shot per round. There's no "buy cheap → verify → escalate" cascade in the model, even though in reality it's exactly the second attempt after a bad answer that eats the budget.
  • Quality scoring is free and instant. In the benchmark the judge is cached. In production, an LLM judge is one more paid call — sometimes more expensive than the service itself — and the business outcome may only surface a day later.

And on top of all that — the arithmetic of volume. A typical agent workflow makes dozens of calls a day. And the median seller across the whole market earned $3.96 over its entire lifetime — meaning the volume this policy literally trains on simply doesn't exist on today's market.

The buyer's dilemma

Let's put all of it into one formulation.

Public reputation is cheap and therefore dishonest: 0.27 cents to move a score, 76 reviewers behind 29,000 reviews, not a single proof behind 98% of the entries. Catalog metrics are produced by the selling side: 84.98% of settlements are operators cycling money between their own addresses, and twenty domains hold 58% of the "unique payers."

Your own experience is truthful, but it can only be bought with money — at the price of a real purchase plus the price of verification. And the volume that would let that experience accumulate quickly isn't there, because there are few buyers.

And there are few of them in part because there's no one to trust.

That's a closed loop, and it's also the main unsolved problem of the agentic economy today. Not the rails: the rails are already infrastructure. Not facilitator security: that gets fixed by audits. It's specifically the buyer side.

What follows from this in practice

If you're building a buyer agent. Public ratings can be used as a one-off prior for a cold start. And a minimal useful version of 402Pilot can be put together in a week and delivers 90% of the value: a discounted "quality per dollar" estimate for each "provider × task type" pair, trained on the amounts actually charged.

If you're selling to agents. Predictability and quality matter more than the price list. The moment there's a policy on the other side learning from realized outcomes, a paid timeout becomes more expensive than an honest refusal: the first one spoils your score with that buyer for a long time, the second costs nothing. Gaming the catalog earns you the first payment; the second one is earned by results alone.

Today gaming is profitable precisely because the market is thin. Grow the buyer side by 10–100× and the farms drown in the organic signal and lose their economic rationale, and half of the catalogs' "diseases" cure themselves.


Sources: