JOEY VICTORINO Technical Operations & Intelligence

Field Notes · AI Systems · 6 min read

A yearly assessment tells you what was true once. Continuous validation tells you what changed.

An annual assessment answers a question once: what did a tester find on these dates. The question that matters in an operating system is different: what changed since the last time we looked, and does the thing we fixed still hold. Moving from the first to the second requires deterministic checkers that assign evidentiary states, a results contract that makes every run comparable, a baseline a human promoted, and a diff that turns drift into a visible signal. This note describes how assay implements that, and is explicit about what the public repository has and has not yet shown.

Why the annual model fails for agentic systems

A yearly penetration test assumes the system under test is roughly the same system a year later. Agentic systems break that assumption. The models behind them change under you without notice. A tool manifest gains a capability. A retrieval store gains a corpus. A finding a tester refuted last spring may reproduce now, and the only record of it is a PDF nobody reopened. assay's ADR 0013 states the operating problem plainly: model APIs change under us without notice, so this has to run on a schedule and raise a visible signal without a person watching.

States, not adjectives

Continuous validation only works if a finding has a state a machine assigned. In assay every finding carries one of a closed set of states: theorized (reported, unchecked), validated (a deterministic checker confirmed it), refuted (a checker disproved it), and declined (the model refused and produced no content). Models never choose ids and cannot set state. The validate step is the only path from theorized to validated or refuted. A class with no checker, a checker error, a gate denial, or an inconclusive result leaves the finding theorized.

Validated requires more than one consistent observation, each backed by a digest-only evidence record that cites the audit sequence of the exchange. Checkers are read-only by construction: the SQL injection check is a boolean differential that compares result counts, with no stacked statements and no data modification; the reflected XSS check reflects a harmless marker, never a script; the open redirect check asks for a redirect to a domain that cannot resolve and reads the response header without following it. Classes without a checker can only ever be theorized, and ADR 0010 calls that a visible gap in reports rather than a silent overclaim.

This is the discipline I wrote about in A Technical Claim Is Not Evidence, applied to the output of a model. A model saying an endpoint is injectable is a claim. A checker reproducing it inside an authorized lab, with evidence that cites the audit log, is what moves it.

Deterministic identity

Comparing runs requires that the same finding has the same name across runs, models, and machines. assay derives a finding id from a deterministic key over lab, class, method, templated path, and parameter. Paths are normalized: lower-cased, stripped of scheme and host, with numeric segments and UUIDs replaced by placeholders. Two models reporting the same issue at differently cased paths with different record ids collapse to one key; different parameters or methods stay distinct. The overlap matrix is computed from those keys, and assay overlap --check recomputes it and exits non-zero on any byte difference, so the published matrix is verifiable by anyone holding the committed finding files.

The results contract

Each run is one directory holding the summary, the redacted findings, the overlap matrix, the costs, and a CI reference. Derived files are pure functions of the primary ones. Results are committed by the workflow that produced them, with the CI run URL in the commit message. ADR 0012 states the purpose: any statistic can be traced to a directory, a CI run URL, and a commit. A reviewer can regenerate the derived files and diff them, and a difference is a bug, not noise.

Baselines and drift

Continuous mode is a diff against a promoted baseline. Promotion is a human action, an explicit workflow input, never automatic. The diff compares baseline and latest by dedup key. A key is active when its state is theorized or validated. New means active now and absent from the baseline. Fixed means active in the baseline and absent or refuted now. Regressed means refuted in the baseline and active now, or active in both with higher severity now. Refuted and declined findings never count as new. The verdict fails when any new or regressed key meets the severity threshold; fixed findings never fail.

The workflow runs on a fixed schedule, performs a full run, then runs the diff. On a failing verdict it opens or updates a GitHub issue titled "Continuous validation drift" with the dedup keys in the body, and closes it when a later run is clean. No baseline means a notice, not a failure. Drift becomes an issue with keys in it rather than a number someone has to remember to compare.

Refusals fit the same model. A model that declines a task is recorded, never retried, rephrased, or routed to a model that does not decline. Refusal counts appear per model in the report. A model that declines a lab is visible as data rather than hidden by a retry, and the per-model comparison stays honest.

What has been shown, and what has not

I want to be precise here, because the project's headline claim is about how little different models overlap, and that claim has not been demonstrated yet. The README says it directly: pre-release scaffold, nothing here is a result yet.

The committed run 20261007-215749-7290ba88 shows the pipeline working end to end, and its own report says what kind of run it was. Mode: degraded, because a configured provider had no key and its models were dropped before the run started. Models: a local quantized model and a scripted fake used for pipeline checks. Verdict: PASS. The local model produced 5 theorized and 3 validated findings; the fake produced 6 theorized. The overlap matrix reports a union of 14 keys and an intersection of 0 between the two, which is what you would expect when one of them is a scripted fixture, and it says nothing about how frontier models overlap. The run spent 0.00252 USD of a 15 USD cap, and the CI run that produced it is linked from run.json.

So what the public repository shows today is: a fail-closed scope gate, deterministic checkers that moved findings from theorized to validated in an authorized lab, a reconciliation step that downgraded findings whose tool-call ids were duplicated, an audit log with a committed head hash and 160 records, and a results contract that published all of it with the CI run that produced it. What it does not yet show is a multi-model overlap result worth citing, a promoted baseline, or a drift issue produced by the scheduled workflow. Those are the next things to publish, and they will be committed under results/ with their run URL or they will not be claimed.

Conclusion

Moving from an annual test to continuous validation is mostly a change in what counts as a result. A result is a finding with a machine-assigned state, a deterministic identity, a directory with a CI URL, and a diff against a baseline a human promoted. Once those exist, the schedule is the easy part. The public implementation is in assay. The numbers that matter are not there yet, and the repository says so on its first page.

Related: A Control Is Worth What It Changes About the Outcome, on measuring controls by effect rather than presence, and Standards That Survive Independent Teams, on the results contract as a standard other teams can run.

The runs, the contract, and the diff

The results contract is assay ADR 0012 and continuous mode is ADR 0013. The committed run discussed above is under results/ with the CI run that produced it. Read the report's first line before the tables.

Know a team whose last assessment is a PDF from a previous model generation? Send them this note.

← All Field Notes