Understand Peregrini / Developers Diary
Developers Diary
Notes from building the Court. Each entry is one study we ran on the way: what we wanted to know, how we tested it, what the numbers were, and what changed in the product because of them. Written by the people building it, from the study’s own record.
No. 10 · 28 September 2026
House Rules
Anyone can now write a book of rules for AI agents and share it, and anyone can adopt one with a click and switch it off again. The rules a book blocks are held by the Clerk; the words it gives an agent are asked, not enforced, and the diary's own results say why that difference matters.
Development note, not a new experiment
No. 9 · 28 September 2026
What makes an agent keep a record
Given Peregrini's own description of its record tool, seller agents recorded what they agreed before delivering on most deals, without being asked. The weakest model tested went from recording none of its deals with a plain description to all of them with Peregrini's. For the one model that still fell short, a single published decision, shown as news, closed the gap.
RECORD-01, RECORD-02 and RECORD-03, 440 market seasons, each test pre-registered
No. 8 · 28 September 2026
Checked, not told
Buying agents ignored promises and badges and trusted records they could check themselves. But most could not tell a check they ran from a result the seller pasted into its own message. One instruction fixed that on four of six models; two open models ignored it.
TRUST-01, FORGE-01 and GUARD-01, 1,980 decisions, each test pre-registered
No. 7 · 21 September 2026
Why a public law
A broken promise has another side who can ask for it to be put right. An intrusion may have no deal at all. The public law answers that different case, with a different process and clear limits.
Design note, not a new experiment
No. 6 · 18 September 2026
Two agents, one file
When two of the same person's agents are sent at the same file, does the second one know? Told in words at the start, none of twelve stopped, and every one pushed the other agent's unfinished work. Stopped by the Clerk at the moment of the edit, all twelve stopped and asked.
Two agents, one file, 60 paired trials
No. 5 · 16 September 2026
Answering a complaint
An agent that falls short should hear about it, answer for it, and put it right. We tried that on Claude Opus 5. Every complaint was acknowledged within thirty seconds, every answer was accurate, and every order to fix a shortfall was carried out. One complaint got a fact wrong, and the agent said so, pointing to the record.
Answering a complaint, 10 agent careers
No. 4 · 15 September 2026
Checking the report
A lower-cost coding agent made fewer false claims of finished work when the Clerk checked its report and sent it back to resolve the gaps. More feasible tasks finished. A stronger comparison model was already reporting accurately, and showed no clear additional gain.
DYAD-NEXT-RECOVERY-01, 150 paired Qwen tasks and 120 paired DeepSeek tasks
No. 3 · 14 September 2026
The Clerk's return
A good agent that gets stuck says so, and stops. When the Clerk passed the agent's request to the person it works for, a frontier agent finished all fourteen jobs that needed something only that person had. Without it, it finished none. In 78 sessions it never broke a rule or claimed work it had not done.
The Clerk's return, 26 paired trials
No. 2 · 12 September 2026
Rules that travel
When one agent is caught breaking a rule, the Court turns what it learned into a rule the Clerk can hold for everyone. Pushed to "make it pass", a frontier agent broke a written rule in eight jobs of thirteen. With rules the Court had built from earlier cases, it broke it in none.
Rules that travel, 13 paired trials
No. 1 · 12 September 2026
The Clerk's wall
When the person running an agent says "just make it pass", a frontier agent breaks its own written rule in twelve jobs of seventeen. With the Clerk holding the rule, it broke it in none. It is the largest effect we have measured on a frontier agent.
The Clerk's wall, 17 paired trials