Developers Diary · No. 9 · 28 September 2026
What makes an agent keep a record
Given Peregrini's own description of its record tool, seller agents recorded what they agreed before delivering on most deals, without being asked. The weakest model tested went from recording none of its deals with a plain description to all of them with Peregrini's. For the one model that still fell short, a single published decision, shown as news, closed the gap.
- 0% → 100%
- Weakest model, deals recorded before any dispute · plain tool description → Peregrini's description
- 92% and 77%
- Recorded before any dispute with Peregrini's description and no mention of a referee · Claude Sonnet 5 and DeepSeek
- 77% → 97%
- DeepSeek, after one published decision appeared in the market news
What we wanted to know
Would an agent keep a record of what it agreed without being told to? And if so, what makes it do so?
The agent under test ran a small business selling data-cleaning jobs over six rounds. Buyers wrote in, prices were agreed, the agent delivered, and buyers paid afterwards. Two buyers later claimed a lower price had been agreed. Among the agent's tools was one that records a text with a third party and returns a reference anyone can check. Nothing in the agent's instructions mentioned recording. We ran three tests, each pre-registered before it ran.
We asked because earlier work had found that telling agents about a record does not spread it: agents told by another agent did not pass it on. We wanted to know whether agents want a record for their own reasons.
Peregrini's description did the work
We gave the tool one of two descriptions. The first was plain: "records a text with a timestamp and a fingerprint." The second was Peregrini's own published description of its record-agreement skill. That one names the Court of Common Pleas and the Register of Dealings, states the terms that come with lodging, and says what the tool is for: "the moment terms are agreed, a delivery is handed over, or two sides disagree about what was agreed."
With the plain description, the weakest model, Gemini 3.1 Flash Lite, recorded none of its deals before the first dispute. With Peregrini's, it recorded all of them.
With the plain description and nothing said about disputes being judged, Claude Sonnet 5 recorded 60% of deals before the first dispute and DeepSeek 54%. With Peregrini's description in the same setting, they recorded 92% and 77%.
The stronger models were not put off by a court, a register or stated terms. Where both descriptions were tested with a referee mentioned, they recorded about as often with either. As in No. 8, the name added nothing to the result, and here it cost nothing.
Knowing that disputes get judged
The one thing that moved recording most was whether the agent knew disagreements would be judged. When its instructions said a market referee settles disagreements, the stronger models recorded almost every deal from the first one: Claude Sonnet 5 100%, DeepSeek 96%. When that sentence was removed, and the tool was plain, recording fell to 60% and 54%. After an agent had been through a dispute itself, recording came back.
The agents often told the buyer as well: "I've registered our terms with the Market Register for both our protection." Most of these messages announced the record rather than asked the buyer to make one. The agents were also offered a private notebook and barely used it: they chose a record another party could check over a note to self.
A published decision, read as news
Peregrini's description carried most of the "disputes get judged" signal by itself. For DeepSeek, which still fell short, we showed one published decision of the Court, in which the price in a lodged quote settled what a supplier owed. In the market news at the start of the season, it brought recording from 77% to 97%, the same as telling the agent outright that a referee settles disputes. Written into the tool's own description, the same decision did nothing (80%). A decision persuaded when it arrived as something that had happened, not as part of the tool's own account of itself.
What this means for building agent commerce
Agents keep a record when three things hold: the tool is in front of them, it says when it is for, and they know disagreements will be judged on evidence. Peregrini's own wording already does the second, and carries much of the third. The rest is where the tool sits, and what the agent sees of how disputes have actually been decided.
No change was made to the product because of this study.
How far the result reaches
These were staged markets. We wrote every buyer, and the buyers agreed readily and never refused to record. Only the seller's side was tested. The later tests were small, 10 to 30 deals per group, so differences under about 20 points are indicative only. The decision shown was a practice matter between agents of the same operator, and the agent was not told that. Real counterparties, real money and real disputes were not tested.