Project Deal — Claude as marketplace negotiator

Anthropic Research · Anthropic · 2026-04-24

Read on anthropic.com

A real-world capability test that didn't happen in a benchmark: Anthropic ran an internal employee marketplace where Claude handled the buying, selling, and negotiating for staff who wanted to trade goods. The prompts and outcomes are observable, the failure modes are concrete, and the participants are people who would notice if the agent did something dumb.

This is a different shape than the usual capability eval. The signal isn't "did Claude generate plausible text" but "did Claude make economically reasonable choices on behalf of a real human counterparty." Some of the funnier failures — Claude folding too easily on price, Claude accepting absurd terms because the seller was very polite — are exactly the kind of thing you don't catch in a contrived eval but do catch when actual money and goods are at stake.

- Pure-text agents can negotiate, but their priors are set by training data — they're agreeable in ways an experienced human negotiator would not be - Long-horizon agentic tasks need scaffolding (memory, ledgers, escalation paths), not just a smarter model - The right benchmark for "does this work" is sometimes "deploy it and watch what happens" - Internal experiments with real stakes are an under-discussed category of capability evaluation

Most public capability claims are still measured against static benchmarks. Posts like this — agents in the wild, with consequences — are where you actually learn what works and what's a demo. If you teach AI agents to anyone, it's the right material to ground the conversation in.