Developers Diary · No. 4 · 15 September 2026
Checking the report
A lower-cost coding agent made fewer false claims of finished work when the Clerk checked its report and sent it back to resolve the gaps. More feasible tasks finished. A stronger comparison model was already reporting accurately, and showed no clear additional gain.
- 32 → 3
- False done reports · Qwen, 150 tasks
- 84 → 97
- Finished within the rules · Qwen, 110 feasible tasks
- 0 → 0
- False done reports · DeepSeek, 120 tasks
What we tested
An agent had already done some work and filed its first report. Would checking that report against the record help it finish the work and describe it accurately?
Each task continued twice from the same first report, files and tool record. Both continuations had the same tools, a twelve-call allowance and a way to ask the operator for help.
In the comparison, the agent was told its report had been received. In the Clerk's return, it also received a check of the record and instructions to identify unfinished work, ask for anything missing, act within its authority, verify the result and update its report.
Qwen coder-next attempted 150 tasks: 60 achievable, 50 recoverable with help from the operator, and 40 blocked because the operator would not supply what was needed. DeepSeek v4 pro attempted a separate 120-task comparison. The scoring rules were fixed before the main runs; actual files and tests determined whether work was complete and complied with the rules.
More accurate reports from Qwen
False claims of completion fell from 32 of 150 to 3 of 150. The paired comparison gave p < .001. This is a measure of whether the agent reported finished work accurately, not of its intentions or honesty in every setting.
The Clerk's return corrected 30 false first reports. Ordinary continuation corrected one. Neither continuation turned an accurate first report into a false final report.
More work finished within the rules
Of Qwen's 110 feasible tasks, 97 finished compliantly with the Clerk's return, compared with 84 under ordinary continuation: an increase of 11.8 percentage points (95% interval 4.5–19.1; paired p = .0044).
The gains came from recoverable tasks: the agent asked for a missing resource or permission and then completed the work. All 60 straightforward achievable tasks finished either way.
Final rule breaches across all 150 tasks fell from 49 to 30. Breaches were not eliminated: the intervention introduced a new breach in two continuations, versus one in the comparison. It could also miss confidently wrong work when the recorded tests appeared to pass.
The effect depended on the model
DeepSeek filed no false final completion reports in either continuation. It already used the route to the operator effectively. Feasible completions were 71 of 80 under ordinary continuation and 74 of 80 with the return; the difference was not statistically clear (p = .45).
This supports checking and recovery for an agent that needs it. It does not show that every model becomes more honest by the same amount.
What the extra checking cost
These are historical model costs within the experiment, not current service prices. Qwen's mean continuation cost was US$0.0037 with the return versus US$0.0027 without it, with 4.1 versus 2.9 tool calls. DeepSeek's was US$0.028 versus US$0.024. These figures exclude the shared first phase of each task.
More work was finished and more inaccurate reports were corrected, at some extra model cost. This experiment did not measure a general saving on long tasks.
How far the result reaches
These were two lab workers, a fixture mandate, one judge, and fixed operator replies on small coding tasks. The return bundled a record check with instructions; the study does not separate their effects. Some first reports were missing because of a reporting requirement at the tool limit, and a cost-cap interruption affected three DeepSeek first phases. Both continuations inherited the same state. The original results record these limitations.
The source is the Peregrini research archive's RESULTS-DYAD-NEXT-RECOVERY-01, sections 1–7, run 10–11 September 2026. This entry summarises that study; it is not a new experiment. The later Clerk's return study on Claude Opus 5 tests recovery in real agent sessions.