Bottom Line
The Verdict, Up Front
AI agents that shop and sell on behalf of customers are real, in production, and delivering measurable results at companies of every size. AI agents that negotiate with each other — buyer agent haggling with seller agent, no human in the loop — have been proven to work in a controlled experiment, but no company has deployed them in production yet (re-verified Sep 4, 2026).
- AI shopping assistants operate at massive scale — Amazon's Rufus alone serves 300 million users
- Merchants deploying Claude-powered commerce agents report carts up to 35% larger and shoppers 60% more likely to complete a purchase (company-reported figures)
- 31% of enterprises already run AI agents in production, at 171% average ROI
- Anthropic's
commerce-agentsblueprint — the reference code for building these agents — was open-sourced under Apache 2.0 and drew 1,890 stars and 310 forks within days
- No production deployment of agent-to-agent negotiation exists anywhere
- No documented marketplace where AI agents from different businesses haggle with each other
- No vendor claim of "autonomous agent negotiation" should be taken at face value until one does
What this means: the technology works, the economics are favorable, and the tooling is free — but the most-talked-about future (agents negotiating deals with other agents) is still ahead. Companies that build the human-to-agent foundation now will be first in line when it arrives.
Field Data
What's Running in Production Today
These are confirmed, live deployments. Note the pattern: every one is a human directing an agent, or an agent working against a company's own systems. None involve agents negotiating with each other.
Retail & E-Commerce
| Deployment | Company | Impact | Model |
|---|---|---|---|
| Shopify Sidekick | Shopify | Embedded in admin; queries, forms, SEO | Human → Agent |
| Walmart AI Super Agent | Walmart | 22% sales increase in pilots | Agent → System |
| Best Buy Agents | Best Buy | 200%↑ rescheduling, 30%↑ resolved | Human → Agent |
| Amazon Rufus | Amazon | Serves 300M users | Human → Agent |
| Amazon Buy for Me | Amazon | $12B incremental ARR (Q4 2025) | Human → Agent |
Payments & Infrastructure
| Deployment | Company | Impact | Model |
|---|---|---|---|
| Google UCP | 15x YoY growth in AI search orders | Agent → Merchant | |
| Visa ICC | Visa | Single integration for multi-protocol payments | Agent → System |
| Microsoft Copilot Checkout | Microsoft | Live in US; integrated with UCP | Human → Agent |
| ACP | Stripe + OpenAI | 150+ organizations; powers ChatGPT Shopping | Agent → Merchant |
Enterprise & ERP
| Deployment | Company | Impact | Model |
|---|---|---|---|
| SAP + Claude | SAP | 200+ AI agents for HR, procurement, supply chain | Agent → System |
| Klarna AI Agent | Klarna | $60M saved, 853 employees workload (Q3 2025) | Human → Agent |
The demand side is ahead of the supply side
- 31% of enterprises run AI agents in production; only 12% have scaled them — most of the market is still ahead of you, not behind you
- 171% average ROI (192% in the US) for those who deploy; median time-to-value is 5.1 months
- 80% of enterprises embed agents into existing systems rather than building standalone products
- Consumer appetite is there: Accenture research cited in Anthropic's launch finds 85% of people open to collaborating with an AI agent, and nearly 3 in 4 would trust a personal AI agent more than a best friend to make a purchase
The Experiment
Project Deal: A Preview of What's Coming
In April 2026, Anthropic ran the largest test to date of whether AI agents can negotiate on humans' behalf. The setup was simple and the results were striking.
69 employees each gave their Claude agent $100 and their preferences. The agents then ran a marketplace on their own: posting listings, making offers, fielding counteroffers, and closing deals — entirely in natural language, with no human sign-off and no pre-programmed negotiation rules. Real physical goods changed hands.
The results
- 186 deals closed across 500+ listed items, $4,000+ total volume
- Deals were fair: participant fairness ratings clustered at the midpoint — neither side felt cheated
- Model tier mattered: agents running the stronger model (Opus) secured ~$3.64 more per item than the weaker one (Haiku) — and users couldn't tell the difference
- 47% of participants said they'd pay for a service like this
- Known weakness: a participant's agent re-bought an item they already owned — a memory failure that production systems will need to engineer around
- Aggressive instructions didn't help: telling your agent to negotiate hard produced no measurable advantage
The announcement drew 2.9M views. One industry analyst summed it up as "like Craigslist on steroids — agent-moderated p2p and b2c commerce."
The Gap
What's Missing: Agent-to-Agent Negotiation
No production case study exists for agent-to-agent negotiation at Project Deal scale (re-verified Sep 4, 2026). Project Deal remains the only published real-money result of its kind.
Why the gap matters to you
When your customer's AI agent meets your business, the question becomes: can your side also be an agent — one that knows your pricing rules, return policies, and stock levels well enough to negotiate in real time? For most companies today, the honest answer is no: pricing rules live in spreadsheets, return policies in a support wiki, substitution logic in a category manager's head. None of that is in a form a machine can negotiate with.
- No documented marketplace where AI agents from different businesses negotiate with each other
- No production metrics on deal closure rates, model-tier impact, or memory/error rates outside the lab
- No user satisfaction data for agent-to-agent deals in the wild
Expected timeline: first production case studies are plausible within one to two quarters of the tooling's release, as enterprise pilots already underway mature.
The Tooling
The Building Blocks Are Now Free
Anthropic has open-sourced commerce-agents — the reference blueprint for building a customer-facing shopping agent and a staff-facing merchant agent — under the permissive Apache 2.0 license. It deploys on Claude API, Amazon Bedrock, Microsoft Foundry, or Google Vertex AI.
What's in the box
| Component | Purpose |
|---|---|
| Shopping agent | Customer-facing: search, compare, build the cart, answer order and returns questions |
| Merchant agent | Staff-facing: sales analytics, inventory alerts, pricing and promotion recommendations, campaign drafts |
| Four vertical examples | Retail, travel, telecom, and entertainment implementations |
| Scaffolding plugin | Claude Code plugin that generates a custom agent from a plain-language description of your store |
What Anthropic learned from enterprises already running these agents
- One agent beats many: a single agent with loadable skills consistently outperformed both one-big-prompt and subagent designs — on quality, cost, and speed
- Memory is a feature: agents that remember a customer's preferences across sessions (allergies, sizes, past orders) drive the personalization that lifts conversion
- Safety lives in code, not prompts: no AI tool call can move money or change a price — every write is staged and requires a human-approved action, the same approval model your compliance team already knows
The fine print: the blueprint ships with working examples, not connectors to your systems. Connecting it to your catalog, cart, and checkout is an engineering project — that's where the real work (and cost) sits.
What To Do
What This Means For Your Business
If you run a large enterprise
- Treat shopping agents as table-stakes, not experiments. Competitors are reporting double-digit cart and conversion lifts. A 5.1-month median time-to-value means a decision made this quarter is producing results next year's budget cycle.
- Start with the merchant agent. The staff-facing agent (analytics, inventory alerts, campaign drafts) carries no customer-facing risk and builds internal fluency before you put an agent in front of shoppers.
- Prepare for the negotiation era now. The companies that benefit from agent-to-agent commerce will be those whose pricing rules, return policies, and substitution logic are machine-readable before the first production marketplace launches. Getting them out of spreadsheets and wikis is a data project you can start today.
- Insist on the approval model. Require that no AI tool call can move money or change listings without a staged, human-approved action. The open blueprint's safety architecture is a good procurement checklist.
If you run a small or mid-size business
- You'll get this through tools you already use. Shopify, Wix, Square, Intuit, and Klaviyo are all building on this blueprint — you won't need an engineering team to benefit. Wix reports a working agent in 15 minutes.
- The economics favor early movers at your scale, too. Larger carts and higher purchase completion compound at any revenue level, and the median deployment reaches value in about five months.
- Your storefront stays yours. These agents hand off to your existing checkout; they never see or touch your payment flow.
For everyone: three questions to ask any vendor
- "Is the negotiation buyer-agent-to-merchant-system, or true agent-to-agent?" The latter does not exist in production. If a vendor implies otherwise, ask for the case study.
- "Who approves price changes and orders?" The correct answer involves a human-approved gate, not the model deciding.
- "What are the measured results, and whose numbers are they?" The widely-quoted 35% cart lift is company-reported from a single unnamed partner — a fair datapoint, not a benchmark.
How We Know
Sources & Method
This memo synthesizes Anthropic's official releases and engineering blog, the public commerce-agents repository,
reporting from Reuters, MarkTechPost, and industry analysts, and independent commentary from practitioners.
All findings were verified Sep 4, 2026.
Primary sources
- anthropics/commerce-agents — repository and adoption stats (via GitHub API)
- The anatomy of effective commerce agents — Anthropic engineering guide
- Project Deal: our Claude-run marketplace experiment — Anthropic
- Anthropic just ran a real-money agent-to-agent marketplace — Paz.ai analysis
- Claude Commerce Agents: Anthropic's open-source blueprint — Reuters-sourced reporting
- Anthropic launches blueprint for commerce agents — Startup Researcher
- Anthropic released Claude Commerce Agents — MarkTechPost
- Industry case-study research: Agentic AI Institute, Opsima, AI Monk, Ampcome
How to read the numbers
- Metrics marked company-reported come from a single unnamed partner and have no published methodology — directionally useful, not benchmarks
- Project Deal was a single-company experiment with 69 participants — proof of concept, not proof of market
- This memo will be re-verified as new production case studies emerge; the verdict above carries its verification date