The model validator reran the overnight Value at Risk calculation on their own workstation using a frozen input package. The scenario losses agreed closely, but the reported percentile differed in the eighth significant digit. That was the entirety of the finding. It sat in the tracking system for three quarters because nobody owned the tolerance, and the compliance policy demanded zero discrepancy. The underlying model never changed.

In binary floating-point arithmetic, addition is not associative. Summing a million returns in a different sequence yields a different total. The processor is doing exactly what the IEEE 754 standard specifies. The variation originates in the shape of the reduction tree.

When a risk engine aggregates a massive array across multiple threads, it does not iterate sequentially from the first index to the last. Thread one sums the first chunk, thread two sums the second, and the sub-totals are combined. Altering the thread count alters the tree, which in turn alters the truncation. A weekend resize of the GPU node pool changes the summation order, moving the final risk number in the ninth digit.

You cannot fix a topology problem with better rounding. Kahan summation reduces the error relative to the exact mathematical sum, but two different reduction trees will still produce two different compensated results. It yields a more accurate answer without providing any additional reproducibility.

The final digit is irrelevant until a business threshold rests exactly on top of it. Smooth profit and loss vectors summed in a different order move in the fifteenth significant digit, which is immaterial. But Value at Risk is an order statistic. A risk limit is a threshold. A barrier option is a comparison. Marginal VaR is a difference of two quantiles.

At each of those boundaries, continuous input noise becomes a discrete output change.

Calculating the 99th percentile across 500 simulated scenarios without interpolation simply selects the fifth-worst loss. If the scenarios are reordered by a fraction of a basis point due to a different parallel reduction, the calculation picks the sixth-worst loss instead. The distance between those two losses has nothing to do with floating-point precision and everything to do with the shape of the portfolio’s loss distribution. For a concentrated book, that gap can be substantial.

A VaR limit set at a round number illustrates the same mechanic. Monday’s run on the legacy on-premises cluster prints just under the limit. Tuesday’s run on the new cloud node pool prints just over it. The market moved two basis points, but the limit breach notification was triggered by the new parallel reduction tree. Defending a risk limit becomes impossible when the limit itself shifts upon recalculation.

This is where the argument about acceptable differences usually stalls, because the tolerance interval has to be non-empty. It must be looser than the spread across your own compute fleet and tighter than the smallest change that triggers a business action.

If those two conditions cannot both hold, the control is broken regardless of the arithmetic.

Setting a tolerance band only after a comparison fails yields an explanation, not a control. The useful tolerance is tied to a decision, not to the number of decimal places on a screen. If a difference of 40 cents on a 1.2 billion book does not alter a capital allocation or a hedging decision, it is irrelevant to the business. But it costs three weeks of two teams proving it is irrelevant, every quarter, forever. The cost of reconciliation routinely exceeds the cost of the error itself.

When setting the tolerance, start from the smallest change that would alter a decision, then round up to something defensible. Check that it comfortably exceeds the observed fleet spread. If it does not, fix the design or move the threshold. Do not tighten the tolerance to a number the hardware cannot hit.

The noise also leaks into downstream processes. Run a hedge optimizer or a capital allocator against a VaR number that carries floating-point jitter, and it will converge to the jitter. The objective’s gradient with respect to the noise is often larger than the gradient with respect to the market. You get hedge ratios that move between consecutive nights with no market move and no position change. Making the noise deterministic does not fix it; the optimizer will reliably find the same spurious optimum every night. You have to make the noise smaller than the effect you care about, or make the objective smooth.

Vendor migrations hit the same wall. Two firms running the same trades through the same standardized margin model will sometimes face a collateral dispute over a few hundred dollars. The root cause is the summation order in the sensitivity aggregation, and neither side can point to a clause that dictates whose order is canonical. If a vendor’s new engine uses Fused Multiply-Add instructions at the hardware level, retaining full precision for intermediate products, it is technically more accurate. But if the legacy calculation is the regulatory baseline, procurement stalls because the business refuses to sign off on a system that produces a different risk number for the same portfolio.

Reproducibility is a design choice, not a byproduct of writing correct code.

The most efficient structural improvement pairs a fixed reduction topology with counter-based random number generation. Tile the scenarios into fixed-size blocks summed in a fixed index order, regardless of which thread produced them. Address random numbers by path and timestep instead of by stream position. Those two changes remove most of the practical variance.

Decouple the decision from the valuation. Determine barrier knockouts from the official fixing the instrument references, not from the simulated path. Evaluate limit checks against a number computed by a fixed, single-threaded reference path. Tying a barrier decision to the simulated path rather than the official fixing turns a rounding error into a full notional exposure.

Keep a single-threaded, pinned-build reference configuration and treat it as the definition of the number. Production runs against it, and the difference between them is a monitored quantity, not an incident. Store the input hash, the build hash, the math library version, the thread count, and the tolerance alongside the baseline value. When a test fails, the first question is which of those changed, and that question should be answerable in a minute.

A reference value lacking a provenance record is nothing more than a rumor.

Regulators and validation teams do not offer leniency for IEEE 754 floating-point constraints. They expect the system to reproduce its own output. When the interval between the achievable hardware spread and the smallest decision-relevant change is empty, the right move is to change the measure or change the threshold. That is a design finding, and it is a much better conversation than the one about whether a difference in the ninth digit is acceptable.