polycratia

Why a VEX document should be diffed claim by claim

· 8 min read

Nobody reads your VEX document twice. A customer reads the first one you publish, and after that they read the difference between the new one and the copy they already hold: which advisories you have newly admitted, which conclusions you changed, which ones you have closed. If your tooling cannot produce that difference, every release asks each downstream reader to re-triage a whole document from scratch, and from 11 September 2026 the EU Cyber Resilience Act expects manufacturers to answer the "are you affected" question quickly, in writing, for every product they ship. Quickly and repeatedly means the delta is the deliverable.

I built vexdesk (https://github.com/polycratia/vexdesk) around that observation. The diff subcommand turned out to be the part that forced every other design decision to be stated precisely.

What a text diff tells you, and why it is the wrong thing#

A VEX document is JSON, so the obvious move is to diff two of them with the tool you already have. That tool reports what the bytes did. It says nothing about what the claims did.

Statement order in a VEX document is a serialisation detail. Add a component to the inventory, regenerate, and a block of statements shifts; a line-oriented diff renders that as deletions followed by insertions, and you have to reconstruct by eye that nothing changed. The changes that genuinely matter are not textually large either. A conclusion that stayed not_affected but now rests on a different justification is a one-token edit buried in a wall of moved lines.

I have had this exact problem in payments, in a form that predates VEX by a long way. Reconciling a bank statement against internal records only works if each movement is keyed (attributed by its requisites) and compared as a set. Position in the file means nothing, the same movement can turn up under a different reference, and the output you want is not "these bytes differ" but "this is the one movement neither side agrees about". A VEX diff has the same shape: keyed comparison, with the unmatched items promoted to the top of the report.

One vulnerability against one product#

So the first decision is the identity of a claim, and in VEX it is not the statement object. It is the pair: one vulnerability against one product. That pair is the key both documents are indexed by, and the rest (ordering, timestamps, how the statements happen to be grouped) is serialisation.

Once comparison is keyed, the vocabulary of changes falls out of it. A key present now and absent before is new. A key present on both sides with a different status is restated. A key present on both sides with the same status but a different justification is rejustified:

console
$ vexdesk diff released/vex-2.0.0.json vex.json
3 change(s), 4 claim(s) unchanged

CHANGE       ADVISORY      PRODUCT                                      DETAIL
new          FIXTURE-0005  pkg:npm/cogwheel@4.1.0                       under_investigation
restated     FIXTURE-0004  pkg:npm/cogwheel@4.1.0                       affected → fixed
rejustified  FIXTURE-0001  pkg:golang/github.com/example/widget@v1.2.3  justification vulnerable_code_not_in_execute_path → component_not_present

Needs attention (1):
  FIXTURE-0005  pkg:npm/cogwheel@4.1.0  under_investigation

Four claims unchanged, and the report says so in one line instead of showing them. Reordered statements do not show up at all. That is more than a cosmetic improvement: a customer stops reading a document when the document makes them look at everything. A keyed diff makes them look at three rows.

A conclusion that held for a new reason is a change#

The rejustified row is the one I would defend hardest, because a status-only comparison drops it on the floor. The status did not move: it was not_affected before and it is not_affected now. Only the justification changed.

That is still a different claim. The justification is the claim's argument, and a reviewer accepted the argument, not the word. Here is what a decision looks like going in:

json
{
  "decisions": [
    {
      "vulnerability": "FIXTURE-0001",
      "product": "pkg:golang/github.com/example/widget@v1.2.3",
      "status": "not_affected",
      "justification": "vulnerable_code_not_in_execute_path",
      "impact_statement": "the affected parser is only reached from the admin importer, which this build does not include"
    }
  ]
}

vulnerable_code_not_in_execute_path is a reachability argument: the component is in the build, the vulnerable function is not reached. component_not_present is a composition argument: the thing is not shipped at all. A security reviewer who validated the first one validated a claim about call paths, probably by reading them. If the justification quietly becomes the second one, their earlier review no longer covers what the document now asserts, even though the verdict is identical. A conclusion that held for a new reason is the thing that has to be re-read, so it gets its own change class and its own row.

That is also the row that makes the diff useful internally and not only to customers. A justification that drifts between releases is usually a sign the decisions file is being edited to make a gate pass, rather than because someone looked again.

An upgrade is one move, and ambiguity is not a place to guess#

Upgrading a dependency breaks pure keyed comparison, because the product is part of the key. Bump cogwheel from 4.0.0 to 4.1.0 and a naive set diff reports a claim vanishing and an unrelated claim arriving. Both halves mislead you: the first reads as a finding you quietly dropped, the second as a problem you just acquired.

So a claim about the same vulnerability on the same package at a different version gets paired, and reported as one move. The restated FIXTURE-0004 ... affected → fixed row above is exactly that. But only when the pairing is unambiguous. If there are two candidates on either side, vexdesk reports them as they stand rather than guessing which one became which.

That restraint is deliberate, and it is the opposite of what a heuristic wants to do. A wrong pairing does not produce a missing row. It produces a confident sentence about a history that did not happen, and the reader has no way to tell it apart from a true one. Two honest unpaired rows cost somebody a minute of reading. One invented pairing costs their trust in every other row in the table. It is the same rule the matcher follows elsewhere in the tool: a component with no package URL, or a range that cannot be ordered, comes out as unknown with a reason attached rather than as a clean result.

The gate is asymmetric#

The last piece is what the exit code means, and the useful answer is not "the documents differ". Differing is normal, every release differs. What a release should stop for is the current document opening work the previous one did not: a claim that is affected or under_investigation now and was not before.

sh
#!/bin/sh
set -eu

vexdesk vex -sbom sbom.cyclonedx.json -advisories ./advisories \
            -decisions decisions.json -author "Example Ltd" -o vex.json

if vexdesk diff released/vex-2.0.0.json vex.json; then
  echo "no new open work since the last release"
else
  echo "gated: the new document opens work the published one did not" >&2
  exit 1
fi

Closing work does not fail the build. A claim that moved from affected to fixed, or from under_investigation to not_affected with a justification, is the pipeline working. Only the direction that adds unreviewed exposure is a stop condition, which is what makes the gate survivable: a check that fires on every change gets disabled within two sprints.

There are two gates in the workflow and they sit at different points. match exits 1 when anything in the comparison against the advisory set needs attention: that is the triage gate, early, before anyone has concluded anything. diff is the publication gate, and it is about the document you are about to hand a customer. Where each one lands in a release script is written out in docs/cra-workflow.md in the repository.

One more asymmetry, easy to get backwards: documents already issued are read leniently, including ones that break rules the tool enforces on its own output. A published document is a fact in someone's hands. Refusing to parse it hides the statements that most need correcting, and showing those is the whole purpose of the diff.

What I would do differently#

I started with diff as a reporting convenience, a nicer view over two files I already had. It is not that. The identity of a claim is the only decision in the tool I could not change later without invalidating every historical comparison a customer has already made, and I arrived at it by writing the writer first and the comparison afterwards. Starting again, I would write the claim key and the change vocabulary before anything emitted JSON, because the document format sits downstream of them.

I would also be slower to add pairing cleverness. Every heuristic I considered beyond "same vulnerability, same package, different version" improved the common case and introduced a way to narrate a history that did not happen. The version that refuses more often is the one I still trust.

The document is the artifact. What I ship on each release is the difference between two of them.

react

$ new-project --brief

or email hey@polycratia.com