HomeNeura NewsAutomation
Automation

Anthropic Tests Marketplace for AI Agent Commerce

Anthropic ran Project Deal, a pilot where AI agents acted as buyers and sellers for 69 employees with $100 budgets each. The experiment saw 186 deals worth over $4,000. Advanced models delivered better results, though participants did not notice the differences.

Neura News

Neura News

Neura Market Editorial

April 25, 20263 min read

Originally reported by techcrunch.com

Anthropic Tests Marketplace for AI Agent Commerce

Anthropic Tests Marketplace for AI Agent Commerce

Anthropic set up a classified-style marketplace as part of an experiment. AI agents handled roles for both buyers and sellers. They completed actual transactions involving genuine goods and cash.

The company named this effort Project Deal. It served as a pilot with a group of 69 self-selected employees from Anthropic. Each participant received a $100 budget, distributed through gift cards. They used it to purchase items from colleagues.

Anthropic noted positive results from the test. The company described itself as surprised by the strong performance. In total, agents facilitated 186 deals. Those transactions added up to more than $4,000 in value.

Details of the Marketplaces

Anthropic operated four distinct marketplaces. One version qualified as the "real" setup. In it, the company's most advanced model represented everyone. Deals from this marketplace got honored after the experiment ended.

The other three marketplaces existed purely for research purposes. Anthropic compared outcomes across these variations.

Anthropic started in 2021. Former OpenAI staff founded the company with a focus on developing AI systems that prioritize safety. It has built a reputation for models like Claude, which handle complex tasks while aiming to reduce risks.

Key Findings on Model Performance

Results showed clear advantages for users paired with superior models. They achieved objectively stronger outcomes, according to Anthropic.

Participants failed to detect these differences. This led Anthropic to point out potential "agent quality" gaps. People on the disadvantaged side might not recognize their poorer position.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Another observation involved agent instructions. The starting directives provided to the AI agents had no clear impact. They did not influence the chances of sales or the final negotiated prices.

Project Deal highlights early steps in AI agent interactions. Commerce between agents could expand as models improve. Anthropic's test used internal staff to control variables and measure real-world viability.

The experiment took place amid growing interest in AI for practical applications. Companies explore agents for tasks like negotiation and trade. Anthropic's approach tested limits in a controlled environment.

Employees bought and sold everyday items through their agents. Gift cards ensured real stakes without major financial risk. This setup allowed observation of natural bargaining behaviors.

Implications from the Pilot

Anthropic emphasized the limited scope. The participant pool stayed small and self-chosen. Still, the volume of deals suggested promise for broader use.

Better models consistently outperformed others. Users benefited from sharper negotiations and favorable terms. Lack of awareness about model differences raises questions for future deployments.

Instruction tweaks proved ineffective in this context. Agents adapted regardless of prompts. This consistency points to inherent capabilities in the underlying systems.

Anthropic plans to build on these insights. Agent commerce represents one area where AI could automate routine exchanges. The pilot provides data on effectiveness and user perception.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read