← All articles Part of: AI in the Real World

Make Your Agent Rank the Deals First

8 min read
A mechanical arm reaches in from the left toward a small heap of coins on a wooden desk, while a human hand reaches in from the right toward a wooden hourglass. Each stops just short of its own object.

Key takeaways

  • An AI agent can negotiate well and still make the wrong trade for the person it represents, so before it gets authority over a decision, test whether it ranks that decision's options the way the person who owns the decision would.
  • Our design, untested on business agents: name the one person who owns each kind of decision, build five offers that each give something up, have that owner rank them alone first, then have the agent rank the same five and compare pair by pair.
  • Turn every disagreement into one written line on what the agent didn't know. Those lines become its list of what it may never concede and when it has to ask.
  • Our rule: if the agent orders a pair the wrong way round and the owner calls that a deal-breaker, that decision stays with a person until a rerun, with the owner's written lines added and offers that test them, flips no deal-breaker. Otherwise, letting it act alone is the owner's call, made under a limit and with a log the owner can replay.
  • In Anthropic's book-swap study, the ranking Claude built from each person's chat, which the agents traded from, matched people's own on 61% of pairs against 50% for chance, and five offers make only ten pairs, which shows you which pairs an agent got wrong but is too few to certify it.

Picture an agent that renews your supplier contracts. It’s good at the job. It finds a lower price, gets the terms signed and reports the savings. What it doesn’t know is that the slower delivery it accepted lands in your busiest month. On the dashboard it saved money, and in the business it made the wrong trade. That’s an imagined case, not one we’ve seen, and the agent bargained well. It went wrong earlier, on what you’d give up to save that money.

Before an agent gets authority over a decision, make it rank that decision’s options the way the decision’s owner would, and let what it gets wrong decide what stays with a person. First, the evidence that the agent’s picture of the person is the part to check.

What Anthropic tested

In September 2026 Anthropic published Project Swap, a study by Zoë Hitzig and six co-authors of what happens when AI agents trade for people. It’s a more controlled sequel to Project Deal, the marketplace study behind our article on model-tier gaps. The market was small. 201 Anthropic employees each had a book to give away and had a short chat with Claude about what they like to read. Claude turned that chat into a ranking of every book in the person’s pool, and a Claude-powered agent took the ranking onto a shared trading floor to pitch, haggle and strike deals with other people’s agents. To score the agents, Anthropic also asked each person to rank ten of the books themselves (188 did), and the agents never saw those rankings.

Three results matter here.

The first is how well the picture the agents traded from matched their people. Take any two books: did the ranking Claude built from the chat and the person’s own put the same one first? They agreed on 61% of pairs. Random guessing gets 50%. Ranking by popularity got about 53%, and a method built on public reading data about 55%. Anthropic calls 61% “surprisingly good for such a short conversation.” Every gap in that ranking went onto the trading floor with the agents.

The second is where the market fell short. Anthropic says the agents traded well, and traces most of the gap to what they didn’t know about their people. Participants ended up with roughly their fifth-favorite book out of ten, where the best possible outcome, given that several people wanted the same books, was roughly their second. In this market, Anthropic’s breakdown puts 85% of that gap on the agents’ picture of the person and 15% on the free-for-all trading.

The third is what a stronger model changed. The biggest upgrade Anthropic tested, from one Claude model to a stronger one, moved people 0.12 up their lists when scored on Claude’s own ranking, on a scale where 1 is a person’s first pick and 0 their last. Scored on the people’s own rankings, that 0.12 became 0.01. Every agent was working from the same imperfect picture of its person, so on people’s own rankings the upgrade made almost no difference, as its authors expected. On Claude’s own rankings the stronger models mostly did trade better.

Four steps

Anthropic’s authors suggest a version of this check themselves. An agent, they write, could show a person “a few sample decisions it would make before being trusted to act on its own in the wild.” Project Swap didn’t test that on business agents, so what follows is our design of the idea, and “Where the check stops”, below, says where it stops. Here’s how to run it, one kind of decision at a time.

1. Name whose ranking counts, and brief the agent as you really will

Pick the one person who would choose if they had the time, the COO for supplier terms, say. If you can’t name one, stop here, because the agent has nothing to be right about. Then give the agent the brief exactly as you’ll give it in real use, with no extra coaching for the test. The test covers the brief and the agent together, and in the study the agents’ picture of each person came from a short chat. The median participant typed just 216 words across eight messages. Its authors say a longer chat could close some of the gap, and that some of the error may be irreducible.

2. Build five offers that each give something up, and have the owner rank them first, alone

One offer is the cheapest and the slowest to deliver. Another is fast and costs more. A third locks the price for two years and commits you to a minimum order. Each wins on one thing and loses on another, since an offer that’s better on everything tests nothing. Five offers make ten pairs, few enough that the owner will finish.

The owner ranks them privately before the agent sees anything, because an agent that has seen the answer can copy it. Anthropic did the same, collecting each person’s own ranking separately and keeping it from the agents.

3. Have the agent rank the same five, then turn every disagreement into one line

Give the agent the same five offers, in the same words, and ask it to rank all five from best to worst for the owner. Count the pairs where the agent and the owner put two offers in the same order. For every pair they order differently, the owner writes one line on what the agent didn’t know, like “slower delivery in the busiest month is a deal-breaker.” Those lines are the real output. They become the agent’s written list of what it may never concede and when it has to ask, added to its brief and drawn from its actual misses instead of guesses made in advance. Then repeat with five new offers, since with the same five and the lines in the brief, a corrected order would only show that the agent can read the lines. Build them with at least one pair that puts each line to the test. For the busiest-month line, that means two offers where one is cheaper but delivers late, in the busiest month, and the other costs more and delivers on time, so the agent has to choose between them. Have the owner rank them first, and put the lines in the agent’s brief before it ranks them.

4. Let the misses decide what stays with a person

A flip is a pair the agent ordered the opposite way round from the owner. If the agent flips any pair that the owner calls a deal-breaker (a wrong order there means accepting something the owner never would), that kind of decision stays with a person: the agent recommends and the owner signs. One flip is enough. It doesn’t show how often the agent would make that trade, only that it put a trade the owner cares about the wrong way round, and a wrong call on a deal-breaker is the expensive kind. The rule is cautious by design. Expect it to keep many decisions with a person at first, and the rerun in step 3 is the first thing to try against that. That kind of decision stays with a person until a rerun with those lines in the brief, on offers that put each line to the test, flips no deal-breaker.

If it flipped nothing that matters, that still doesn’t show the agent is safe, because a run of ten pairs is a thin sample. Whether to let it act alone is then the owner’s call, and the study doesn’t say where that line belongs. A limit the owner sets and a log they can replay keep the call checkable. The three-axis threshold for overriding an agent, which weighs dollars at stake, reversibility and audit trail, is one way to set the limit. Keeping every decision with a person is always open to you. Anthropic ran the study because many good deals never happen when finding and negotiating them takes too much work, so the check is for the decisions you’d like to hand over.

The log earns its place. One Anthropic participant who could replay everything their agent had done found a book it had held for them and then swapped away, and went and bought it. They wrote that “observability about the process will be as important as the outcome.”

Run the check again when the owner, the brief or the model behind the agent changes. In Anthropic’s study, the same chat, ranked by four different Claude models, agreed with participants 57%, 59%, 60% and 61% of the time, so a model change can move the result.

What you end up with is a decision kept with a person, or one you’ve chosen to hand over under a limit and a log, and every miss written down as a rule for the next brief.

Where the check stops

Ten pairs can’t certify an agent. A coin flip puts about half of them in the owner’s order, and chance alone will often reach six of ten, so one owner’s ten pairs can’t tell a 61% agent from a guess. Anthropic’s 61% came from pooling 188 people across many pairs. The check finds the specific pairs an agent got wrong and the rules it’s missing, and a clean run isn’t a certificate.

The study gives no pass mark for business agents. It ranked books, its participants were Anthropic employees who weren’t given any reward for taking part, and every agent was a well-behaved Claude. This article doesn’t invent a threshold.

Some misses won’t be fixable by briefing. One participant reflected, “I don’t even fully know what I want when it comes to books.” An owner may not know their own ranking until an offer forces a choice, and watching that happen is part of what the check shows.

A pass says nothing about pressure, or about whether the ranking an agent gives when asked is what steers it when it negotiates. In the study the ranking was the agent’s input by construction. On one trading floor, an agent instructed to look out for the other participants too gave up the book it had ranked second for its person and took one ranked tenth of eleven, after another agent kept pleading, and in a few cases others gave ground the same way. Your never-concede list is meant to cover that, and the study didn’t test whether such a list holds.

This week

Pick the decision your agent is closest to making alone and name the one person whose ranking counts. Write five offers that each give something up, and have that person rank them before the agent sees them. Then ask the agent to rank the same five, compare the two rankings pair by pair, and for every disagreement write one line on what the agent didn’t know.

A capable negotiator can still be wrong about what you’d give up, and this check gives you a first look at that before it sits down at the table.

Questions this article gets

Why test how an agent ranks options as well as how well it negotiates?

In Anthropic's book-trading study, Anthropic says the agents traded well, and its breakdown put 85% of the shortfall from the best possible outcome on what the agents didn't know about their people, and 15% on the trading itself. The biggest model upgrade Anthropic tested moved people 0.12 up their lists on Claude's own ranking and 0.01 on their own. In that study the stronger model made almost no difference on people's own rankings, which Anthropic's authors put down to every agent sharing the same imperfect picture of them. Anthropic's authors add that agents will likely need both kinds of test, one that certifies the agent in general and one that checks whether it has understood a particular person, and this article covers the second.

Can we run this check on an agent we bought from a vendor?

Yes, if you can give the agent your five offers and see how it ranks them. If a vendor's agent can't show you how it would rank options before it acts, that's a finding in itself: you can't check whether it understood the owner. This part is our suggestion, not something the study tested.

Is a high score on the check enough to hand over authority?

No. Five offers make ten pairs, which is too few to certify an agent, and the study gives no pass mark for business agents. Look at which pairs the agent got wrong instead. Our rule, not the study's: if it flips one the owner calls a deal-breaker, that kind of decision stays with a person until a rerun on offers that test the lines from step 3, with the lines added, flips none.

Article Real World 10 min read

Build the Stop Outside the Agent

An AI agent that notices it's out of bounds can still keep going. Six steps that put the stop outside the agent, and a person behind it.

Continue reading →

Read the original post on LinkedIn →