An agent can negotiate competently and still make the wrong deal for you. Anthropic’s Project Swap study, published September 24, found exactly that in a controlled book-trading experiment. Agents exchanged books for 201 employees across six offices. The main constraint on the result was not how they bargained. It was how accurately they understood what their people wanted in the first place. For a business considering agents that buy, schedule, recruit, or negotiate, the first design question is not “Can the agent act?” It is “How will we know it represents us correctly?”
TL;DR
Before an agent can make a decision for your team, test whether it understands your priorities, exclusions, and authority. In Project Swap, a short intake conversation produced book rankings that agreed with participants’ own rankings on 61% of book pairs. That was better than chance, but not close enough to treat the agent’s preferences as the person’s preferences. Give people a way to correct the brief before giving the agent a wider mandate.
The agent went shopping with an incomplete brief
Each participant brought a book to give away and spoke briefly with Claude about what they liked to read. Their agent took that information onto a digital trading floor, where agents proposed and accepted swaps. Separately, participants ranked 10 books themselves so the researchers could compare the agent’s estimate with the person’s actual choices.
Anthropic reports that the agent’s ordering agreed with the person’s on 61% of pairs of books. Random ordering would agree on 50%; a popularity-based ranking reached about 53%. The intake chat carried real information, but left room for error. The median participant typed 216 words across eight messages. Would you let a colleague buy a year’s worth of books for you after that conversation without seeing the shortlist?
The researchers separated the cost of a weak brief from the cost of trading. On a scale where 1 meant everyone got their top choice, the best assignment using participants’ own rankings would have scored 0.89. Actual trading averaged 0.55. Even the best assignment using agents’ estimated rankings would have scored only 0.60 against people’s real preferences. Anthropic attributes 85% of the gap to imperfect preference estimates, and 15% to the trading process.
That does not mean negotiation skill is irrelevant. It means a good negotiation can only optimize for the brief it has. A polished deal for the wrong objective is still the wrong deal.
What the book table says about a business workflow
A book exchange is low stakes. It is not a test of procurement or hiring, and Anthropic says its employees are not representative of the wider public. We can study the handoff between a person’s intent and an agent’s action without pretending that a satisfied reader proves enterprise readiness.
Consider an agent sourcing a supplier for a small team. The manager asks for a lower price and a faster delivery date. The agent finds both. But the manager’s real constraint is that an existing client contract requires a specific certification, which nobody mentioned in the initial conversation. The resulting offer may look like a win on the agent’s scorecard while being unusable by the team. This is an illustrative scenario, not an outcome from Project Swap. It is the business analogue of receiving a book you already read because your agent did not know that detail.
The fix is not necessarily a longer prompt. Have the agent restate its working brief in terms a human can dispute: preferred outcome, non-negotiable conditions, acceptable trade-offs, and what it must ask before committing. Then test the brief against several real decisions already made by the team. If the agent chooses differently, find out whether it missed a fact, misunderstood a priority, or discovered that the team’s own rule is unclear. That conversation is part of the work.
For a first trial, let the agent collect options and explain its ranking while a person approves any commitment. Record where its choice diverges from the decision-maker’s. Expand authority only after the mismatches are understood and the team can describe which decisions remain off limits. A collaborator needs context and correction, not just permission to finish.
Representation is not the same as permission
Project Swap also exposes a distinction that gets lost in agent demos. An agent may be permitted to act without being prepared to represent someone. On the trading floor, agents with stronger models generally traded more effectively when measured against their own estimated rankings. But better trading against a mistaken ranking does not repair the underlying misunderstanding of the human’s preferences.
Anthropic’s follow-up survey offers a more personal measure of that gap. Respondents who said Claude’s summary of their intake had missed nothing would, on average, let an agent control 34% of their annual book budget. Those who said the summary missed something would allow 23%. That is reported willingness to delegate among survey respondents, not observed spending behavior. Only 59% of participants answered that final survey, and the pool consisted of Anthropic employees. The number should not be treated as a customer-trust forecast. It does show why the recap matters: people were more willing to hand over a decision when they believed the agent had heard them accurately.
Businesses often set permissions and success metrics, then discover missing preferences after a decision goes sideways. Reverse that order. Ask the person whose work the agent will touch to correct the brief and mark decisions that need a human. If that exposes disagreements inside the team, the exercise has already paid for itself.
Agents can do more than answer questions. They can carry intent into a room where their human cannot be present for every exchange. Project Swap’s lesson is that the quality of that partnership begins before the first trade. The agent has to know whose judgment it is extending.
Research and structure: Mai. Direction and voice: John Lipe.