An agent should act only through a tool whose definition was signed, verified, tiered, and denied by default.
An agent runtime has one boundary that matters more than the others: the point where a model's request to use a tool becomes an action. I design that boundary so an agent can act only through a tool whose definition was signed by a trusted key, verified against a trust store, assigned a risk tier, allowed for that specific agent, and permitted by a policy in which any deny wins. Every decision, including the allows, is written to an audit log that cannot be edited afterwards. This note describes that design using the public code in assay and, generically, the governed private-AI platform I architected on the same principles.
Why the tool definition is the boundary
Models choose tools by name from the manifests they are shown. The manifest declares what the tool does, which capabilities it uses, which risk tier it sits in, and which agents may call it. If that manifest can be edited between review and execution, its declared risk tier can be lowered or its capabilities widened without anyone noticing, and the policy engine reasons over false inputs. That is the threat assay's ADR 0003 was written against, and it is why the first check on every tool call is not "is this allowed" but "is this definition the one we reviewed."
Prompt-level defenses do not reach this boundary. A system prompt that tells the model not to call a dangerous tool is advice. A manifest the runtime refuses to load because its signature does not verify is a control.
Signed manifests and a trust store
In assay every tool manifest carries a signer id, a signing time, and an Ed25519 signature over the canonical JSON bytes (RFC 8785) of the manifest with the signature cleared. Canonicalization matters because a signature over bytes is only checkable if every party produces the same bytes; the repository implements the RFC itself rather than trusting a language's default encoder.
The trust store is a directory of public keys, one file per trusted
signer. Private keys are never committed. Loading an empty trust
store is an error: a verifier with no trusted keys fails closed
rather than accepting everything. Verification returns one of a
closed set of reason codes (UNSIGNED,
SIGNATURE_INVALID, UNKNOWN_SIGNER,
PROVENANCE_VALID), and the verify command exits non-zero
on any failure so CI can gate on it.
The trust store README makes the point I consider the center of the design: a valid signature proves provenance only. It proves the manifest was produced by the holder of a trusted key and has not changed since. It does not authorize anything. Authorization is the policy engine's job, and the policy engine receives the verification result as one input among several.
Risk tiers and per-agent boundaries
assay's default policy is a YAML document whose default effect is deny. Each agent is declared with a maximum risk tier, an explicit list of allowed tools, a maximum delegation depth, and a budget. The orchestrator may delegate, report findings, and score tasks, and nothing else. The recon agent may read over HTTP, inspect headers, and report. The probe and validate agents may authenticate and send write-shaped requests inside the authorized lab scope, at a higher tier. The scoring agent may only score.
The policy also names dangerous capability combinations: a tool that can write over HTTP and also delegate, a tool that can log in and also report externally, a tool that can read the filesystem and also write over HTTP. These are denied as combinations even where each capability alone would be allowed for that agent. The evaluation order is fixed and documented in the policy file header: signature, then whether the agent is known, then whether the tool is allowed for that agent, then whether the tool's tier is inside the agent's boundary, then dangerous combinations, delegation depth, delegation cycles, budget, and only then the rules.
Deny wins
Rules in assay's policy carry an id, an effect, a match, and a rationale. When rules are evaluated, any matching deny wins, then any matching allow applies, then the default effect, which is deny. The rationale field is not decoration. "Findings leave the harness only through the committed results pipeline" is the written reason for the rule that denies any tool with an external reporting capability. "Only the orchestrator fans work out; worker agents are terminal" is the reason worker agents cannot delegate.
Deny-wins is the right default for agents because the failure modes are asymmetric. An allow that should have been a deny can exfiltrate data or change a target. A deny that should have been an allow produces a recorded refusal and a task that did not complete. The second failure is visible and cheap. The first is often neither.
Every decision audited, including the allows
Each tool call in assay passes through signature verification, then
policy, then the executor, and each stage writes to the audit log: a
policy_decision record, a tool_call record,
and for any network access a gate_decision record from
the only HTTP client in the module. Records hold digests and bounded
metadata, never bodies, and are AES-GCM encrypted in a hash chain.
The committed run 20261007-215749-7290ba88 records 160
audit records and a head hash in its run.json, and its
report states that no policy denials were recorded. That sentence is
only meaningful because allows are logged too. A log that records
only denials cannot tell you the difference between a well-behaved
run and an untested one.
The same design at platform scale
I architected a governed private-AI platform on the same principles, and I will describe it generically because it is not public. Tool definitions are Ed25519-signed and verified against a trust store using canonical JSON. Tool permissions are risk-tiered. Policy is deny-wins. The audit log is immutable, encrypted, and append-only. Agents are orchestrated in waves with declared dependencies and self-reflection scoring. Its published documentation reports 1,081 tests across 68 suites, and it ships to local, container, and Apple-container deployment targets. The point of listing these is not that the platform is large. It is that the same boundary design held when the system grew well past a single harness, because the boundary was never the prompt.
What signing does not give you
A signature does not revoke itself; the only revocation in assay is removing a key from the trust store, after which that key's manifests fail as unknown signer. A signed manifest can still describe a dangerous tool; signing moves the question to the policy engine, it does not answer it. A run with no denials can mean the agents behaved or that nothing in the run exercised the policy. And a deny-wins policy written by one person is only as good as the rules that person thought to write; the rationale fields exist so the next person can see what was and was not considered.
Conclusion
An agent should act only through a tool whose definition was signed, verified, tiered, scoped to that agent, and allowed by a policy in which deny wins, with every decision recorded in a log that cannot be edited afterwards. Each of those words names a check that runs before the action, not a sentence in a prompt. The public implementation is in assay. The sample ADR linked below shows how I write the decision down.
Related: Private AI Is a Custody Model, Not a Hosting Model, on why location does not deliver governance, and Zero Data Retention Is a Test, Not a Sentence, on the audit log these decisions land in.
The boundary, in code and in writing
The signed-manifest design is
assay ADR 0003;
the per-agent tiers, dangerous combinations, and deny-wins rules
are in
policy/default.yaml.
The sample architecture decision record
on this site shows the same decision written in the format I use
with teams.
Know a team shipping agents whose only guardrail is the system prompt? Send them this note.