JOEY VICTORINO Technical Operations & Intelligence

Field Notes · Security Architecture · 6 min read

Zero data retention is a property you test, not a sentence you write.

A sentence in a README that says transcripts are not kept is a promise. A test that plants a random marker in a transcript and then reads every byte the system wrote, ciphertext included, to prove the marker is absent is a property. Retention and auditability are architecture properties, which means they are decided by what fields exist, which code is allowed to open files, and whether the log can be altered after the fact. This note walks through how assay enforces zero data retention and tamper-evident audit, and how ORBIT applies the same stance to agent transcripts from the forensic side.

The sentence and the test

assay's README has the line: zero data retention is a tested property, not a README sentence. ADR 0005 explains why the sentence is not enough. The harness sends prompts to models and receives transcripts that may quote target responses, credentials observed during probing, or source code from the lab. A promise not to keep them is not verifiable. The property has to be enforced by the code structure and checked by tests that fail when it is violated.

Several layers enforce it. First, the sink. The only production implementation of the transcript sink discards its input. A retaining sink exists only behind a debug build tag, so a release binary cannot be configured to keep transcripts. Second, the record shape. The audit log stores digests and bounded metadata, and its writer rejects any string over a fixed length or matching a credential pattern. Pasting a transcript into an audit record is impossible, not discouraged. Third, the tests. A canary test generates a random marker, plants it in a transcript, then scans every regular file under the run's output directory byte by byte, ciphertext included, and asserts the marker is absent. A separate test walks the source tree and fails if any package outside a short allow-list calls a file-creating function.

The consequence the ADR records is the one I care about: retaining a transcript requires adding code in a restricted package and defeating the canary test, both of which are visible in review. The redaction patterns are defense in depth. The primary control is that bodies have no field to live in.

Digests, not bodies

What the audit log does keep is enough to reconstruct a run without reproducing its contents. Each model call is recorded with a prompt hash, a response hash, the claimed tool-call ids and argument digests, tokens, latency, and cost in micro-units. Each outbound HTTP exchange is recorded as a gate decision and a digest. Validation evidence cites the audit sequence number of the exchange that produced it. The committed run exports its decrypted audit records alongside the results, and the results README describes that file as digests and bounded metadata, no bodies.

This is the design choice that reconciles two obligations usually in tension. The operator who wants to know what the system did gets a complete record. The data owner who does not want their content copied into a log gets a log with no content in it.

A log you cannot quietly edit

Auditability fails if the record can be changed after the fact. ADR 0004 describes the structure. Each run has one append-only log file. Each line after the header carries a sequence number, the previous line's hash, a nonce, the ciphertext, and its own hash computed over the previous hash, the nonce, and the ciphertext. The ciphertext is AES-256-GCM over the canonical bytes of the record, with the sequence number, previous hash, and run id as additional authenticated data, so a record cannot be moved to another position or another run. The key is derived per run from a master key that lives only in an environment variable. Reopening an existing file verifies the whole chain first and refuses to continue a broken one.

Verification works with or without the key. Without it, a verifier checks the header, every line hash, and the chain, and detects nonce reuse. With it, the verifier also decrypts each record and checks the inner sequence, previous hash, and run id against the line. Errors name the kind of failure and the line number.

The ADR also records what the chain cannot detect: removing the last line is not detectable from the file alone. The run report therefore commits the head hash and record count. The committed run 20261007-215749-7290ba88 carries both in its run.json: 160 records and the head hash, which lets anyone holding a later export check it against the published summary.

Reconciling what was claimed against what ran

Zero retention solves the data problem. It does not solve the truth problem, because a model's transcript is a claim about what it did, not a record. assay's audit log is written from both sides. The transcript side records what the model claimed to call. The control-plane side records what the executor actually ran. The reconcile step joins them by tool-call id and downgrades any finding whose evidence the control plane does not support.

The committed run shows this in operation on a small scale. Its run.json records reconciliation reasons of duplicate transcript tool-call id and duplicate control tool-call id, 3 each, and the report states that findings whose tool calls appear in that table were downgraded to theorized. The scripted fake model in that run reused identifiers; the harness noticed and lowered its confidence in the affected findings rather than taking the transcript's word.

ORBIT, my agent incident reconstruction lab, is the same stance applied from the forensic side. Its README states the principle: an agent-authored transcript is a claim, not automatically an authoritative record of what executed. It reconciles transcript claims to independent control-plane execution records, detects duplicate tool-call and message identifiers instead of silently overwriting them, hashes each source with SHA-256 before parsing, and keeps every join and finding deterministic. No language model participates in extraction, joins, hashing, or findings. ORBIT is a synthetic work sample and says so; it demonstrates the mechanics, not an investigation of any real incident.

What this does not claim

assay measures and tests its own retention. It does not claim anything about what a model provider keeps; that is governed by the provider's agreement with the operator, and the project's security policy says so. The canary test proves the marker was absent from the files the harness wrote; it does not prove anything about memory or about systems outside the run directory. The hash chain proves the log was not altered after writing; it does not prove the writer was honest. Reconciliation catches disagreement between two records; it cannot catch a lie that both records tell together, which is why the control-plane record needs to be written by a component the agent cannot reach.

Conclusion

Retention and auditability are properties of structure. A log with no field for bodies cannot leak bodies. A chain with authenticated position cannot be reordered without detection. A test that plants a marker and scans every byte turns a promise into a failing build. And a transcript reconciled against the control plane becomes evidence only where the two agree. Write the sentence if you like. Then write the test.

Related: Evidence Gaps Are Findings, on what a missing record means in an investigation, and Forensic Parsers Should Fail Closed, on tools whose output must be a truthful statement of what they did.

The log, the test, and the reconciliation

The audit log design is assay ADR 0004 and the retention design is ADR 0005. The forensic side, reconciling agent transcripts against control-plane records, is ORBIT. Both are open source and both state their own limits.

Know a team whose retention policy lives in a README? Send them this note.

← All Field Notes