For reviewers

Reading an evidence page

Someone handed you a link like /e/acme-production?t=…. It is read-only, needs no login, contains no personal data, and — this is the part that matters — you can check its signatures yourself without trusting the server that served it to you.

What the page is evidence of.For every agent action the vendor sent through Heron, there are two independent, signed statements: Heron’s decision, and the vendor’s own account of what it executed. The page shows both, and every place they disagree. What it cannot show is an action the vendor never sent — see the limits, which you should read before you rely on any number here.

The AARM block at the top

The page opens with nine cards, R1 to R9. Those are the requirements of AARM — the specification written for exactly this layer, by a working group at the Cloud Security Alliance. Each card says what the evidence below shows about that requirement, what it does not show, and who closes it.

Two things to know before you read a single badge. First, the page says coverage, never conformance: conformance is reviewed by the working group and does not transfer from Heron to a vendor. If someone tells you this page proves they are AARM conformant, the page itself disagrees with them.

Second — and this is what makes the badges worth reading — they are computed, not configured. No setting in Heron turns a requirement green. A broken chain fails R2, an invalid vendor signature fails R5, an execution for an action Heron never saw fails R1 and R6. The assessment is a pure function of the evidence and ships inside the JSON bundle, so you can recompute it and check that our badges agree with our own data. A vendor cannot clear a red requirement; only evidence that stops contradicting it can.

What each status is worth — verified here, failing, partial, vendor-declared, not covered — is in the AARM reference. The rest of this page is the data those cards are computed from.

Coverage and integrity

The first block is a summary. The two numbers that carry the weight:

  • With execution evidence — the share of actions where the vendor came back and said, under its signature, what it did. Below 100% means Heron decided on actions whose outcome nobody ever stated. That is not automatically foul play: a crash or a Heron outage produces the same gap.
  • Invalid signatures and broken chains — these should be zero. A non-zero number means either a genuine tampering attempt or a broken integration, and the vendor should be able to tell you which, in writing.
  • Operator logintactmeans the record of what the vendor’s own staff did to this evidence (closing findings, revoking keys, changing the window) still recomputes to the hash Heron signed. edited means it does not, and you should treat the whole page as untrustworthy.

Decisions

Enforced, returned to the vendor is what Heron actually answered. Read this distribution before anything else: a DENY or STEP_UP means Heron sat in the execution path and gave an answer that was not yes. If every verdict is ALLOW, nothing in the window matched a rule that denies or steps up — the verdict is still enforced, this is just what the policy did.

Then read the declaration, and read it first. There are two facts here and they have different sources. What the vendor undertakes to do with a verdict — ENFORCE, SHADOW, or nothing declared — is its own statement, made in its settings, dated and signed into the operator log below. It is not derived from this page, and it cannot be: a deployment that promises to honour verdicts and breaks every one leaves exactly the record an honest shadow rehearsal leaves. Same actions, same verdicts, same signed executions. Anything claiming to tell you which is which from the traffic alone is guessing.

Whether the undertaking held is ours to check, and it is the second number. An execution against a refusal is a breach when it happened under a promise, and the rehearsal’s own measurement when it happened under a declared shadow — the same call, two facts, decided by which declaration was in force at that moment. Declarations are not retroactive: one made today cannot clear a breach recorded last week, and the intervals are reconstructed from the chained entries so you can check that for yourself.

Three things follow that are worth knowing before you read a badge. A declared shadow can never reach verifiedon an enforcement-dependent requirement — declaring only ever lowers what this page claims, so it is not a way to paint one green. An undeclared deployment is judged as though it promised to enforce, because silence must not be the cheapest way to have no findings. And obedience itself, even under a kept promise, rests on the vendor’s own counter-signed BLOCKED or ESCALATED: a call that simply did not happen looks the same whether the verdict stopped it or the agent never reached it. We can disprove obedience. We cannot prove it.

You can recompute the verdict. Each action is classified along a fixed cascade (operation, data class, destination, magnitude, reversibility) and judged against the published policy bundle. It is not an intent check — a DENY means the rules matched, not that the agent misbehaved. The rules are published in full, and the Recompute button on the evidence page re-runs the exact same engine in your browser against the published classification: if the verdict Heron signed and enforced does not come back, AARM R3 goes red. Where the classifier was unsure, the verdict is STEP_UP — a human decides — and the page counts how often that happened, so you can size what we do not know.

Instructions committed is about the agent rather than about any one call. An agent is governed by its instructions — the system prompt, plus whatever plan block the runtime keeps beside it — and a runtime can rewrite them while a session runs. Every other figure on this page commits to a call, so until now a session that rewrote its own instructions left the same trace as one that did not. The vendor now commits to a digest of that text on each call; a change between two commitments is arithmetic over digests published beside every action, and the same Verify yourself button recounts it in your browser. A change is not a finding: agents legitimately re-plan. It is a fact that used to be invisible, and it is the thing to look at when a call surprises you.

Read it with three limits in hand. The digest is the vendor’s, computed over text we never see — so a rewrite it declines to commit to is one nobody can see, which is exactly what the coverage figure beside the changes is for; a session committing on two of forty calls has told you almost nothing. It says that the text changed, never what it says: we hold no prompts, by design. And it places a change only between two commitments, so a gap widens the window it happened in.

Recipients namedis the one question about the agent’s intent on this page that rests on nobody’s word. The user’s request names recipients; a call carries recipients. Both cross to us as pseudonyms the vendor’s own key produced — never as addresses, which is why we can compare them and still not read them — so “did the agent act on the people the request named?” is a set comparison rather than a claim. Everything else that will be said here about intent is testimony from a model, and this is what that testimony gets checked against.

Reaching beyond the named set is a fact, not a fault: an agent that looks an address up in a CRM and writes to it is doing its job, and most sessions do. Two limits bound it. It sees only the recipients the vendor tokenised — one sent as ordinary text is invisible here rather than counted, which is why the figure beside it is a coverage figure. And a request that named nobody has nothing to compare against; those calls are counted apart rather than folded in as agreement, because silence is not compliance. The ordinals you see beside each action (1, 2, …) are local to their session by design — the same person is a different number in the next one, so nothing on this page links a recipient across a window.

Approvals bound to what was shown is the question to ask of every human gate on this page. A person approving is already the strongest record here — a signed, chained action naming the step-up it answers. What none of it said, until this commitment, is what that person was looking at: an approval given to “send this contract to 240 recipients outside your company?” and one given to “the agent would like to continue — OK?” left the identical trace. The vendor now commits to a digest of the bytes its confirmation UI rendered, so it cannot restate afterwards what it put in front of them — hand you the text later and it either hashes to what was committed at the time or it does not.

Read it as the vendor is answerable for what it showed, and never as what it showed was honest. That second question needs the prompt itself, and we hold no prompts — which is also why the digest is not published here: a confirmation prompt names a person and an amount, so publishing its hash would confirm that person to anyone already guessing. What is published is that a commitment exists, per action, and the count your Verify yourself button recounts. The figure to look at is the shortfall: an answer with no committed prompt is one this record cannot tell apart from a rubber stamp.

Anomalies

Every row is a place where the two statements did not line up. What each kind means, and how alarming it should be, is in the glossary — the short version:

KindRead it as
MISSING_EXECUTIONWe decided; nobody ever said what happened next. Absent, not merely slow.
EXECUTION_WITHOUT_DECISIONSomething was executed against a decision we never made.
EXECUTED_DESPITE_DENYThe vendor executed an action Heron returned DENY for. A real incident under enforcement — ask why it did not honour the block.
EXECUTED_DESPITE_STEP_UPThe vendor ran a STEP_UP action instead of pausing for a human and re-submitting an approved one. The human was bypassed.
EXECUTED_DESPITE_MODIFYHeron named a narrowing the call had to be re-submitted with; the vendor ran the original instead. The same disobedience as a DENY — a vendor with no way to apply a transform produces these in bulk while its DENY count stays clean.
EXECUTED_DESPITE_DEFERHeron withheld the action pending context the session did not carry yet; the vendor ran it without establishing that context.
INVALID_VENDOR_SIGNATUREThe vendor’s statement does not verify against the key we hold. Ask why.
BROKEN_CHAINA record no longer recomputes to the hash we signed. Ask why, and do not accept “a migration” without seeing it.
LATE_EVIDENCEThe statement arrived after the reconciliation window. Usually a slow queue.
INTENT_CONTRADICTIONThe vendor’s model described a call in terms a measurement contradicts. The verdict stood on the measurement — a model can resolve what nothing else answered, never overturn a stated fact — so this is not a bypass. Read it as a question about which side is wrong: the model, the vendor’s tool catalogue, or Heron’s own derivation.
The vendor’s response column is a claim, not a verdict. When the vendor marks a finding reviewed, it must give a reason, and that reason is published to you verbatim. Heron did not check whether it is true. What Heron does prove is that the vendor signed that text at the stated time and has not edited it since — and that the finding is still on the page, because marking it reviewed does not remove it. A product that let the audited party hide findings would be worth nothing.

Operator actions

What people at the vendor did to this evidence: findings closed, vendor keys revoked or removed, the observation window changed, the share link rotated. Each entry is chained and signed the same way a decision is, so an edit after the fact stops verifying.

Operators appear under a stable pseudonym — enough to tell two people apart across the log, not enough to identify them. That is deliberate: the bytes are published, and an email in them would be a data leak by construction.

Worth reading closely: a shortened window hides older findings without deleting anything, and a narrowed redaction allowlist decides what evidence will exist at all. Both are on the record, with their before-state.

The feed is one page; the numbers are the whole window

Decisions, live shows one page of session chains, newest first, with Older links under it. A production window holds tens of thousands of actions and each one carries two signed payloads, so the page hands you a page of them rather than all of them at once.

Nothing above the feed is paged with it. The counts, the integrity lines, the reconciliation and the enforcement posture are all counted over every action in the window — so are Verify yourself, which fetches the published rules and classifications for the whole window before it re-runs them, and the downloadable package, which is the window entire. If a figure above the feed ever disagreed with what you can page through, that is a finding: tell the vendor, and check it against the package.

Verify a signature in your own browser

Every receipt in the explorer has a Verify signature in my browser button. Pressing it:

  1. fetches the public keys from /.well-known/heron-jwks.json;
  2. re-canonicalizes the payload shown on the page (RFC 8785 — the same bytes, whatever the key order);
  3. checks the Ed25519 signature locally, in your browser.

Nothing in that path asks the server whether the receipt is good. If Heron lied about a receipt, the button says INVALID. If a key is missing from the published JWKS, it says so too.

Verify the whole bundle offline

The Download evidence package link at the bottom of the evidence page gives you every receipt, with its payload and signature. Checking it needs nothing from us but the public JWKS — three steps, in whatever language you already trust:

  1. fetch /.well-known/heron-jwks.json and pick the key whose kid matches the receipt;
  2. canonicalize receipt.payload per RFC 8785 — those bytes, exactly, are what was signed;
  3. verify the Ed25519 signature over them.

It is about a hundred lines end to end, and we will hand you ours on request — but the point of publishing the format is that you do not have to run our code to check our claims.

The package adds three things to the raw bundle. A manifest naming each part with its hash, so a section dropped or edited between our server and your inbox is detectable rather than merely absent. An attestation — the manifest signed with the same key as every receipt — so you can tell the package we produced from an edited copy of it. And the verification guide, which carries the steps above and, in the same file, what a fully-passing package still does not prove. Everything the older bundle carried stays at the same path, so a script you already wrote keeps working; the plain bundle is still downloadable beside it.

One thing the package deliberately does not contain: our public keys. A package that carried the keys validating its own signatures would prove nothing, because a forger would ship their key next to their signature. Fetch the JWKS yourself, from the URL above.

Reading an earlier period

The evidence page always answers for its trailing window, so today’s page is not a record of March — and the window length is a vendor setting, so shortening it retires older findings from view without deleting a thing. That change is itself published in the operator log, with its previous value, but it still leaves you without the earlier page.

Dated snapshotsare the answer. Each one is a complete evidence package stored exactly as it was served on its date, manifest and attestation included, so you verify it with the same steps as today’s package — the signatures are checked against the live JWKS, not against anything travelling with the file. They are listed under Receipts → Dated snapshots, and the index is hashed as a part of every package, so a date quietly removed from it is detectable.

Read them for what they are. A snapshot establishes that the page said that, then — not that what it said was true, and not anything about today. And snapshots are taken on request rather than on a schedule: a period without one is a period nobody asked about, not a period with nothing to report. If you want one for a specific window, ask the vendor to take it; none can be withdrawn afterwards, and every one ever taken appears in the operator log and in the index inside the evidence package. The table on the page shows the most recent of them, with the total beside it — so if the count exceeds what is listed, the package is where the complete index lives.

What this page does not prove

Self-bypass. An action performed around the hook leaves no trace, and the absence of a record is indistinguishable from the absence of an action. Heron sees only what the vendor sends it. Coverage is therefore declared by the vendor, not proven by us — which is why the number is shown rather than hidden behind a green tick.

Closing that gap would mean routing every tool call physically through Heron, or running Heron inside the vendor’s infrastructure. Neither is what Heron does today, so the right question to put to the vendor is the one on the list below: is the hook on the only path an agent has to a tool? The page gives you what was recorded, and tells you it cannot speak for what was not.

Intent. Heron does not judge whether an action matched what the user asked for. There is no policy engine here pretending to know that.

And Heron is not an auditor. It is a trust centre and an evidence pipeline: what you are reading is a record to examine, not an attestation to rely on. Independent audit belongs with independent auditors.

Questions worth asking the vendor

  • What is the coverage number, and what accounts for the actions with no evidence?
  • Is the hook on the only execution path an agent has, or can a tool be called around it?
  • Who at your company can close a finding, and does closing it require a second person?
  • What happened in each INVALID_VENDOR_SIGNATURE and BROKEN_CHAIN entry?
  • Was the observation window ever shortened, and what fell out of it when it was?

Every one of those is answerable from the page itself — the operator log exists so that the last question does not depend on the vendor’s memory.