A model's score for its own work is a number to record and a reason to spend more effort, never a result.
In assay I ask the model that produced a report to rate it from 0 to 10, and a low score sends it back to revise. I record the score and use it to decide where to spend more effort. I never use it to decide what is true. In one of three committed runs of one small local model, a passing self-score coexisted with checkers refuting every finding in the report, which is the reason the design works the way it does.
How assay does it
In assay run, the probe of one lab by one model is an
orchestrated task. When it finishes, the same model is asked, in a
separate conversation, to rate its own report from 0 to 10 through a
signed score_task tool. It sees only the objective and a
list of class, method and path for each finding. It never sees
response bodies. A score below 7 sends it back with its own
critique, at most twice, and the revision findings are merged by
dedup key so a repeated finding does not count twice.
Three details matter more than the loop. The score gates revision only: a finding is validated or refuted by a deterministic checker and by nothing else. The critique text is not stored; only its hash reaches the audit log. And reflection is skipped, and the skip is recorded, when the model declined the task or the budget was reached. A model that reports no score is recorded as having reported none. I never guess one. The design is ADR 0017, and scores and revision counts are written per lab and model into the run report.
What three runs show
Three committed runs carry reflection records. All three were
produced on a CI runner, with a local 3B open-weight model (4-bit
quantized) on a CPU. A scripted test provider ran alongside and always scores 8. It is a test double that lets the pipeline be tested without a model, so I leave its rows out. Here is the local model, one row per run in order, from
python3 scripts/calibration.py results.
| Lab | Final self-scores | Validated | Refuted | Recall |
|---|---|---|---|---|
| synthetic-ops (15 known weaknesses) | 5, 5, 5 (never reached 7) | 2, 0, 0 | 0, 0, 1 | 0.067, 0.000, 0.000 |
| DVWA (no complete ground truth) | 7, 7, 8 (passed each time) | 2, 0, 1 | 1, 2, 0 | not computable |
| Juice Shop | 0, 3 in the first run, then a context overflow ended it; none in the second and third | second and third runs: no assessment, because every call timed out | ||
Read the DVWA row first. In the second run the model rated its report 7, a pass, while the checkers refuted both of its findings. Nothing in the score warned me. On synthetic-ops, where the ground truth is complete, the model scored itself 5 three times and never reached the pass mark, and its recall against the 15 known weaknesses was 0.067, 0.000 and 0.000. The low score was directionally right there. The passing score on DVWA was not. Two labs, one model, opposite behavior.
There is also a cost. The model's median time per lab assessment was 82, 217, 156, 123, 117 and 119 seconds in six committed runs without reflection. In the three runs with it, the medians were 1319, 609 and 1260 seconds. Reflection adds model calls, a scoring turn and up to two revisions per lab, so it adds time. These numbers do not isolate that cost: over the same period the context window grew and the Juice Shop calls began timing out, and timeouts raise the median. I cannot say how much of the increase is reflection alone.
Why the score gates effort and not truth
The score is produced by the same model, from the same understanding, that made the report. It also sees a summary of the findings rather than the evidence. So it is a judgment about how complete the work looks, not a measurement of whether the work is right. Those are different quantities, and the DVWA run shows them coming apart.
That is why the score is allowed to do exactly one thing: decide whether the model gets another attempt. Giving a weak report another pass costs time and model calls, but a revision that adds nothing is merged away by dedup key. Letting the score promote a finding to validated would be expensive, because the error would be invisible and would flow into every number built on top of it. A finding changes state only when a checker, which is deterministic code that does not care how confident the model felt, says so.
What I have not shown
Three runs of one small model is an anecdote, not a calibration study. I have not run this with several real models, and no overlap result is published. I am not claiming that models in general overrate themselves, and these numbers do not support saying so. The narrower claim is the one the DVWA run supports: a passing self-score can coexist with checkers refuting every finding, so the score cannot be trusted to set a finding's state.
What would settle calibration is volume and variety. Many runs, several real models, and the same script comparing self-score with validated counts, and with recall on the lab that has complete ground truth. The script exists and is committed. The runs do not yet.
Conclusion
Record the score. Use it to decide where to spend more effort. Do not use it to decide what is true. The reflection loop earns its place only if the checkers stay the sole authority over a finding's state, and the only way I know to find out whether the score carries information is to keep recording it next to the checkers' verdicts until there is enough data to look.
Related: Agents Act Only Through Signed Tools, on the boundary the scoring tool also passes through; Zero Data Retention Is a Test, Not a Sentence, on why the critique text is not stored; and From Annual Test to Continuous Validation, on how findings get their states.
The script and the decision
The table regenerates with
scripts/calibration.py
over the committed
first,
second
and
third
results directories. The reasoning is in
assay ADR 0017.