A cryptography inventory has to name the module, not just the line
Most of the cryptography your program runs was written by somebody else. A scan that stops at first-party code will faithfully report the SHA-256 on line 137 of a file you wrote and miss the MD5 inside a transitive module you have never opened, and the second finding is usually the one that decides how long a post-quantum migration actually takes. So an inventory owes you two things past finding things: attribute every finding to the module that introduced it, and say out loud which modules it could not read.
I have been building and running payment and crypto systems since 2018, and the shape of this problem is familiar from reconciliation work. A row you cannot attribute is not a row you can act on. A row that is missing entirely is worse, because nothing in the report separates it from a clean result.
A finding in my code is an edit; a finding in a dependency is a version#
Two different verbs, and conflating them is what makes an inventory useless as a plan.
Cryptography in my own repository is something I change: open the file, change the parameter, run the tests, deploy. The blast radius is known, I am the reviewer, and the calendar cost is a sprint at most.
Cryptography in a dependency is something I have to upgrade away from. The question stops being "what should this be" and becomes "has upstream shipped a release that removed it, and if not, do I fork, replace, or drop the dependency entirely". Each of those has its own calendar cost and its own owner, and each carries a risk of breaking something unrelated. An RSA key size in my own signing code is a refactor. The same algorithm inside a vendored client library is a negotiation with somebody else's roadmap.
That is why attribution is not decoration in the output. Posture tells you severity, the module tells you what the unit of work is. And the unit of work is a module, not a line: one version bump can clear twenty findings at once, or none of them, and a flat list of file positions will not tell you which.
In cbomscope, the module travels with the finding all the way out: into the WHERE column of the table, into -json as module, and into the CycloneDX document as a cbomscope:module property. The coordinate is the identifier everywhere outside go.mod itself, so that is what gets written:
// Module is one dependency named in go.mod.
type Module struct {
Path string
Version string
Indirect bool
// Dir is where the module's source was found: the module cache, a vendor
// directory, or the directory a replace directive points at. Empty when the
// source is not on disk.
Dir string
// Replaced is the coordinate the module was required as, when a replace
// directive redirected it.
Replaced string
}
// Coordinate is path@version, the way a module is named everywhere outside
// go.mod itself.
func (m Module) Coordinate() string {
if m.Version == "" {
return m.Path
}
return m.Path + "@" + m.Version
}The position of a dependency finding is written as the coordinate plus the file inside it, example.com/legacy@v1.4.0/checksum.go:31, so you reach the source from the finding alone, with no guessing about which module cache directory that path came from.
Indirect dependencies run the same code as direct ones#
The requirement list gets read whole, direct and indirect alike. An indirect module does not get a pass in this problem, because it executes in the same process as everything else. The one thing "indirect" tells you is who you have to talk to about it.
func Deps(root string) (DepScan, error) {
modules, err := Modules(root)
if err != nil {
return DepScan{}, err
}
var result DepScan
for _, m := range modules {
if m.Dir == "" {
result.Missing = append(result.Missing, m)
continue
}
found, err := moduleSource(m)
if err != nil {
return DepScan{}, fmt.Errorf("scan %s: %w", m.Coordinate(), err)
}
result.Assets = append(result.Assets, found...)
result.Read = append(result.Read, m)
}
return result, nil
}Look at what the loop does with a module whose source is not on disk. It does not skip it. It moves the module to a different list, and that list is part of the result type instead of a log line.
Somebody else's tree is read on different terms than mine#
When the scanner walks my own repository, test files are fair game: they are code I own and change. When it walks a dependency, they are not. A consumer does not link a dependency's tests, and counting them would inflate the inventory with cryptography that does not ship, which is probably the quickest way to produce a migration plan nobody trusts, because the first three findings someone checks turn out to be in a fixture.
if d.IsDir() {
if path == m.Dir {
return nil
}
name := d.Name()
if name == "vendor" || name == "testdata" ||
strings.HasPrefix(name, ".") || strings.HasPrefix(name, "_") {
return fs.SkipDir
}
return nil
}
if filepath.Ext(path) != ".go" || strings.HasSuffix(path, "_test.go") {
return nil
}
assets, err := File(path)
if err != nil {
return nil
}That last return nil is the other concession: a file that does not parse is skipped rather than failing the whole run. Third-party trees carry files behind build tags for toolchain versions you do not have, generated code, artifacts that were never meant to compile on your machine... Taking the inventory down because one of them will not parse trades a partial answer for no answer.
A silent skip is the failure mode that needs a counterpart in the output, and that is the next section.
go.mod is syntax, not a command#
The requirement list is parsed as text, not handed to the go command. That is a deliberate constraint: an inventory can then be taken with no toolchain and no network, which is what you want in a container that has neither, and what you need when the question is "what is in the thing that shipped" rather than "what does a fresh resolve produce today".
The cost of not asking the toolchain: you have to locate the source yourself, and you honour replace on your own.
func locate(m Module, vendored, cache string) string {
if vendored != "" {
// A vendor directory is the build's answer; the cache is not consulted
// beside it, or the inventory would describe code that is not compiled.
if dir := filepath.Join(vendored, filepath.FromSlash(m.Path)); isDir(dir) {
return dir
}
return ""
}
if cache == "" || m.Version == "" {
return ""
}
dir := filepath.Join(cache, filepath.FromSlash(escape(m.Path))+"@"+escape(m.Version))
if isDir(dir) {
return dir
}
return ""
}Two decisions worth stating plainly. First, when a vendor/ directory exists it is the build's answer and the module cache is not consulted beside it, since otherwise the inventory describes code that is not compiled, which is a subtly wrong document rather than an incomplete one. Second, a replace directive is followed to what actually ships, because that is the code that compiles, and the coordinate it was required as stays in Replaced so the report can still name both. A pinned fork with the MD5 still in it is not cleared by the version number in the require line.
There is also a small piece of file-system trivia that has to be reimplemented: the module cache spells an uppercase letter as ! followed by its lowercase, so a case-insensitive file system cannot collide two module paths. Get that wrong and every module with a capital letter in its path silently becomes "not on disk", which drops you straight into the next problem.
An empty result and an unread module look the same from outside#
This is the part I care about most, and I learned it somewhere else entirely: a parser that returns an empty list when it fails is more dangerous than one that raises, because an empty list is also a legitimate answer. Nothing downstream can separate the two, so the failure propagates as data.
A dependency scan has exactly that shape. "This module contains no cryptography" and "this module's source is not on disk" come out as the same absence of rows. So the second one is not an absence, it is a row:
$ cbomscope scan . -deps
POSTURE ASSET WHERE WHY
broken MD5 example.com/legacy@v1.4.0/checksum.go:31 collisions are practical; unusable for signatures or integrity
quantum_reduced SHA-256 internal/cbom/cbom.go:137 256-bit digest: quantum collision search reduces the margin; SHA-384 or larger is the usual answer
2 asset(s): broken=1 quantum_reduced=1
1 module(s) required by go.mod were not read; run `go mod download` to include them:
example.com/gone@v0.1.0The comment on the field says it better than the prose around it can: an inventory that silently skips a module reads exactly like one that checked it and found nothing. Missing is carried out of the walk in the result type for that reason, next to Read, so a caller cannot get the findings without also being handed the coverage.
The practical consequence is that the interesting number in a dependency scan is not the finding count. It is the ratio of modules read to modules required. A report that says "no broken cryptography in dependencies" after reading nine modules out of sixty is not a clean bill of health, and the only thing standing between you and that mistake is whether the tool bothered to say so.
What I would do differently#
Two things are rough, and both come down to the difference between "present in the tree" and "present in the binary".
A module's whole tree is read, not the package graph reachable from your build. That over-reports: a file behind a build tag that never applies to your platform still contributes findings. I would rather over-report and name the module than under-report, since a named module is cheap to dismiss and a missing one is invisible, but if you are building a plan from this output, read "present" as "needs a look", not "compiled in". Narrowing to the reachable package set is the next step, and it costs exactly the property that makes the current design useful: no toolchain, no network.
The second is coverage as a machine-readable signal. Right now unread modules are named in the human output. If I want CI to refuse a document that covers a fraction of the tree, the count has to be something a pipeline can branch on, not something a person notices. Coverage belongs in the artifact, not only in the terminal.
Neither of those changes the thesis. An inventory exists to tell you which work is yours and which work belongs to somebody else, and it does that by naming the module behind every finding, including the modules it never managed to open.