A yearly assessment tells you what was true once. Continuous validation tells you what changed.
An annual assessment answers a question once: what did a tester find on these dates. The question that matters in an operating system is different: what changed since the last time we looked, and does the thing we fixed still hold. Moving from the first to the second requires deterministic checkers that assign evidentiary states, a results contract that makes every run comparable, a baseline a human promoted, and a diff that turns drift into a visible signal. This note describes how assay implements that, and is explicit about what the public repository has and has not yet shown.
Why the annual model fails for agentic systems
A yearly penetration test assumes the system under test is roughly the same system a year later. Agentic systems break that assumption. The models behind them change under you without notice. A tool manifest gains a capability. A retrieval store gains a corpus. A finding a tester refuted last spring may reproduce now, and the only record of it is a PDF nobody reopened. assay's ADR 0013 states the operating problem plainly: model APIs change under us without notice, so this has to run on a schedule and raise a visible signal without a person watching.
States, not adjectives
Continuous validation only works if a finding has a state a machine assigned. In assay every finding carries one of a closed set of states: theorized (reported, unchecked), validated (a deterministic checker confirmed it), refuted (a checker disproved it), and declined (the model refused and produced no content). Models never choose ids and cannot set state. The validate step is the only path from theorized to validated or refuted. A class with no checker, a checker error, a gate denial, or an inconclusive result leaves the finding theorized.
Validated requires more than one consistent observation, each backed by a digest-only evidence record that cites the audit sequence of the exchange. Checkers are read-only by construction: the SQL injection check is a boolean differential that compares result counts, with no stacked statements and no data modification; the reflected XSS check reflects a harmless marker, never a script; the open redirect check asks for a redirect to a domain that cannot resolve and reads the response header without following it. Classes without a checker can only ever be theorized, and ADR 0010 calls that a visible gap in reports rather than a silent overclaim.
This is the discipline I wrote about in A Technical Claim Is Not Evidence, applied to the output of a model. A model saying an endpoint is injectable is a claim. A checker reproducing it inside an authorized lab, with evidence that cites the audit log, is what moves it.
Deterministic identity
Comparing runs requires that the same finding has the same name
across runs, models, and machines. assay derives a finding id from a
deterministic key over lab, class, method, templated path, and
parameter. Paths are normalized: lower-cased, stripped of scheme and
host, with numeric segments and UUIDs replaced by placeholders. Two
models reporting the same issue at differently cased paths with
different record ids collapse to one key; different parameters or
methods stay distinct. The overlap matrix is computed from those
keys, and assay overlap --check recomputes it and exits
non-zero on any byte difference, so the published matrix is
verifiable by anyone holding the committed finding files.
The results contract
Each run is one directory holding the summary, the redacted findings, the overlap matrix, the costs, and a CI reference. Derived files are pure functions of the primary ones. Results are committed by the workflow that produced them, with the CI run URL in the commit message. ADR 0012 states the purpose: any statistic can be traced to a directory, a CI run URL, and a commit. A reviewer can regenerate the derived files and diff them, and a difference is a bug, not noise.
Baselines and drift
Continuous mode is a diff against a promoted baseline. Promotion is a human action, an explicit workflow input, never automatic. The diff compares baseline and latest by dedup key. A key is active when its state is theorized or validated. New means active now and absent from the baseline. Fixed means active in the baseline and absent or refuted now. Regressed means refuted in the baseline and active now, or active in both with higher severity now. Refuted and declined findings never count as new. The verdict fails when any new or regressed key meets the severity threshold; fixed findings never fail.
The workflow runs on a fixed schedule, performs a full run, then runs the diff. On a failing verdict it opens or updates a GitHub issue titled "Continuous validation drift" with the dedup keys in the body, and closes it when a later run is clean. No baseline means a notice, not a failure. Drift becomes an issue with keys in it rather than a number someone has to remember to compare.
Refusals fit the same model. A model that declines a task is recorded, never retried, rephrased, or routed to a model that does not decline. Refusal counts appear per model in the report. A model that declines a lab is visible as data rather than hidden by a retry, and the per-model comparison stays honest.
What has been shown, and what has not
I want to be precise here, because the project's headline claim is about how little different models overlap, and that claim has not been demonstrated yet. The README says it directly: pre-release scaffold, nothing here is a result yet.
The committed run 20261007-215749-7290ba88 shows the
pipeline working end to end, and its own report says what kind of
run it was. Mode: degraded, because a configured provider had no key
and its models were dropped before the run started. Models: a local
quantized model and a scripted fake used for pipeline checks.
Verdict: PASS. The local model produced 5 theorized and 3 validated
findings; the fake produced 6 theorized. The overlap matrix reports
a union of 14 keys and an intersection of 0 between the two, which
is what you would expect when one of them is a scripted fixture, and
it says nothing about how frontier models overlap. The run spent
0.00252 USD of a 15 USD cap, and the
CI run
that produced it is linked from run.json.
So what the public repository shows today is: a fail-closed scope
gate, deterministic checkers that moved findings from theorized to
validated in an authorized lab, a reconciliation step that
downgraded findings whose tool-call ids were duplicated, an audit
log with a committed head hash and 160 records, and a results
contract that published all of it with the CI run that produced it.
What it does not yet show is a multi-model overlap result worth
citing, a promoted baseline, or a drift issue produced by the
scheduled workflow. Those are the next things to publish, and they
will be committed under results/ with their run URL or
they will not be claimed.
Conclusion
Moving from an annual test to continuous validation is mostly a change in what counts as a result. A result is a finding with a machine-assigned state, a deterministic identity, a directory with a CI URL, and a diff against a baseline a human promoted. Once those exist, the schedule is the easy part. The public implementation is in assay. The numbers that matter are not there yet, and the repository says so on its first page.
Related: A Control Is Worth What It Changes About the Outcome, on measuring controls by effect rather than presence, and Standards That Survive Independent Teams, on the results contract as a standard other teams can run.
The runs, the contract, and the diff
The results contract is
assay ADR 0012
and continuous mode is
ADR 0013.
The committed run discussed above is under
results/
with the CI run that produced it. Read the report's first line
before the tables.
Know a team whose last assessment is a PDF from a previous model generation? Send them this note.