polycratia

The mask is a secret too: value barriers in constant-time C++

· 10 min read

A secret-dependent choice written as mask arithmetic instead of an if is not yet constant time. If the optimizer can work out what the mask holds at the point you use it, it is allowed to specialise the loop back into a branch around a memcpy, which puts back exactly the leak the mask removed. You want two properties here, not one: the mask has to be opaque to the compiler where it is used, and every byte of the destination has to be written whichever way the mask goes.

I have built custodial wallet services and on-chain payment flows where a tag comparison is the only thing standing between a request and somebody else's money. In that code this is the failure I look for most often, because it hides in files that already look fixed.

The original bug, and the fix that is only half of one#

The starting point is the one everybody knows:

cpp
if (std::memcmp(mac, expected, 16) == 0) {
    std::memcpy(session_key, derived_key, 32);
}

Two leaks, not one. memcmp returns as soon as two bytes differ, so timing the call tells an attacker how many leading bytes they guessed right, and the tag falls one byte at a time. That is the famous half. The other half is the copy: whether thirty-two bytes of store traffic happened at all is observable from outside through duration, cache state and page behaviour, and it is observable because it is conditional on the secret answer.

So you fix the comparison. In ctsafe that is equals, which reads every byte and decides afterwards:

cpp
// Reads all 16 bytes. The copy below still branches on the answer.
if (ctsafe::equals(expected_tag, presented_tag)) {
    std::memcpy(session_key, derived_key, 32);
}

Better, and not done. The comparison no longer leaks how far the attacker got, but the acceptance decision still steers control flow, and the work on each side of that branch is not the same work.

The rewrite you are quietly asking the compiler for#

The standard move is to turn the choice into arithmetic. Written by hand it usually comes out close to this, and this is the version not to ship:

cpp
// Not ctsafe. The shape most people write first.
unsigned char m = accepted ? 0xFF : 0x00;
for (std::size_t i = 0; i < n; ++i)
    dst[i] = (dst[i] & ~m) | (src[i] & m);

Read it as a compiler. m has exactly two reachable values, and its provenance is a bool sitting a few statements up the same basic block. For m == 0xFF the loop is dst[i] = src[i]; for m == 0x00 it is dst[i] = dst[i], which stores nothing anybody can observe. Both specialisations are cheaper than the general form, so a good optimizer emits them: a branch on m, a memcpy on one side, nothing on the other. The source has no branch. The binary has the branch back, keyed on the secret, with the store traffic conditional again.

There is no bug to report here. Timing is not in the C++ object model: observable behaviour is I/O and volatile access, and the transformation preserves all of it. Every argument of the form "but then it leaks" is an argument about something the language does not model. The move left to you is to make the rewrite impossible rather than undesirable.

The barrier belongs on the mask#

So the mask stops being a local unsigned char whose history is in plain view and becomes a value the compiler has lost track of. In ctsafe it is a type, produced from a bool by arithmetic rather than by a branch, and consumed by operations that take the mask instead of the question:

cpp
ctsafe::mask accepted = ctsafe::mask_from_bool(ctsafe::equals(expected_tag, presented_tag));
ctsafe::copy_if(accepted, session_key, derived_key, 32);
ctsafe::select(accepted, out, derived_key, fallback_key, 32);

Two separate guarantees are in play, and they get enforced in two different places.

The first is the one people remember: in equals, no branch depends on the contents, and the accumulated difference passes through a compiler barrier so the loop cannot be short-circuited.

The second is the one that gets dropped: the mask itself passes through a value barrier before it is used, so the compiler cannot discover what it holds and rewrite the loop as a branch around a memcpy. Without this, the first barrier's work gets undone at the call site, and the call site is usually in a different file, written by somebody who believed the hard part was already handled.

The distinction in the name is worth keeping. This is not a memory barrier; nothing needs ordering. It is a barrier on a value: the number stays in a register, and the compiler forgets where it came from. At the instruction level it typically costs nothing at all. What it truncates is the dataflow the optimizer is allowed to reason over, which is the entire point, because constant time is a property of the code the optimizer is not allowed to produce.

Write every destination byte either way#

The barrier makes the branch unprofitable to guess at. What makes the two paths actually cost the same is the discipline inside the operation: select and copy_if read both sides and write every byte of the destination whichever way the mask goes, so a rejected copy costs the same stores as an accepted one.

That sentence contains a price, and you want it stated out loud rather than discovered in a profile. Constant time over a buffer means paying the accepted case every time. There is no fast path for rejection, because a fast path for rejection is the leak. For a thirty-two-byte session key nobody will ever notice. If you find yourself wanting this over something large, the design question is why a secret is deciding the fate of something large, not how to make the rejection cheaper.

A few consequences fall out of that shape, and the tests pin them down rather than leaving them as prose. A conditional copy of zero bytes is allowed and does nothing surprising. select may write into one of its own inputs, so the in-place case (choose between what you have and a fallback, into the buffer you already hold) is a supported call rather than a trick. And because the entry point takes a ctsafe::mask, the only way to spend the answer is to let it decide bytes: you cannot accidentally hand a bool to something that is going to branch on it.

Checking that it survived your flags#

A claim about generated code that is only checked by reading source is a claim about last year's compiler. The repository puts this one on a clock, in a dudect-style harness: time one operation over two classes of input, draw the class at random before each measurement, crop the slow tail at many percentiles, then apply Welch's t-test to what is left, failing at |t| > 10. The case for this property is as blunt as the property:

cpp
timing::result copy_if_mask_set_vs_clear(const timing::config& cfg) {
    std::mt19937_64 setup(cfg.seed);
    std::vector<std::uint8_t> source(buffer_bytes);
    std::vector<std::uint8_t> destination(buffer_bytes);
    fill_random(setup, source);
    fill_random(setup, destination);
    ctsafe::mask copy = 0;

    auto prepare = [&copy](std::mt19937_64&, int cls) { copy = ctsafe::mask_from_bool(cls == 1); };
    auto run = [&copy, &source, &destination] {
        ctsafe::copy_if(copy, destination.data(), source.data(), destination.size());
        sink = static_cast<std::uint8_t>(sink | destination[0]);
    };
    return timing::measure("copy_if: the mask set vs clear", cfg, prepare, run);
}

Its expectation is no_difference. Alongside it, control_memcmp is registered with the expectation difference: a known leak the harness is required to find, so a silent run means the measurement still works rather than that the measurement stopped working. The suite runs at -O0, -O2 and -O3, because a property that only holds at -O0 is precisely the bug. And the runner ends by printing the honest version of its own verdict: a clean run found no difference on this machine, which is not the same as there being none.

The erasure side of the same family of bug, memset on a dead buffer, deleted because nothing reads it afterwards, gets checked the other way: in the disassembly, against three fixtures that differ only in how they clear the key. A baseline with no clear at all, the ctsafe::erase version, and the naive memset written on purpose as the control the check has to catch being deleted.

What I would do differently#

For a long time I treated the comparison as the unit of work. Write a constant-time equals, barrier the accumulator, move on. That is how the leak ends up back in the binary one call frame away: the fixed comparison returns a perfectly honest bool, and the next line branches on it.

The rule I apply now is that a secret-derived value and the use of that value are two separate audits. Anywhere such a value reaches a control structure, a length, or an index, it needs to be opaque at that point, not opaque somewhere upstream. In review that turns into a concrete question: for this if, where did the condition come from, and how many bytes of store traffic differ between its two sides?

The other thing I would do earlier is keep a control in the timing suite from the first commit rather than adding it once the suite was green. A harness that finds nothing looks identical whether the code is clean or the measurement is broken, and the only thing telling those apart is a leak of known size that it is obliged to catch.

Finally, the bytes you conditionally copied still have to not outlive the scope that holds them, and that is a different duty with the same adversary: the optimizer removing work nobody observes. A guard answers it in the destructor:

cpp
{
    ctsafe::secret_span secret(key);
    sign(message, secret.data(), secret.size());
    if (ctsafe::equals(secret, presented)) { /* accept */ }
}   // key is zero here

The erase belongs to exactly one span: moving it hands the duty over, erase_now() clears early, and release() hands the duty back to the caller. Comparison is still available, spelled as what it is, since == on a secret is memcmp underneath, so it is deleted rather than merely discouraged.

None of this is clever code. It is the ordinary code, written in the one form the compiler is not permitted to improve.

react

$ new-project --brief

or email hey@polycratia.com