Most tools show you a backtest and call it proof. We hold ourselves to a harder standard, because a business makes real decisions on these numbers. This page says exactly how they were produced, and what they can and cannot promise.
Every figure SettlePoint computes is walk-forward: at any decision moment, the engine may only use information that existed before that moment. A safeguard in the code refuses any calculation that tries to read a price at or after the decision time: it is mechanically impossible for tomorrow’s rate to leak into today’s number.
The three execution approaches were built and calibrated on one span of history, then frozen: their rules locked, byte for byte. Only then were they graded, once, on 51,517 invoices from periods the engine had never seen during design (including 2024–2026) with the test criteria written down before the answer was read, and the single read of the sealed data recorded in an access log. The benchmark for every invoice is the same: what settling immediately would have delivered, in cents on the rate and rand on the amount.
The Defensive approach cut the scatter of due-date settling to roughly a quarter: the certainty it sells, measured
The worst 1-in-20 invoice under the approaches, versus leaving invoices to the deadline with no approach at all
Armed levels reached before the due date in comparable invoices: a measured base rate, quoted with its sample, never a promise
Testing didn’t stop at launch. Every read the desk serves is logged the day it’s shown, append-only. No one can rewrite what was said. Its core claims are then graded against what the market actually did next, and the running scorecard is open inside the app. When a claim’s accuracy slips below its recorded standard, the desk stops serving that claim rather than quietly serving a worse one.
And the bank margin check is not a model claim at all: it is verifiable arithmetic against the live fair forward. Check it yourself.
The comparison you see in a replay isn’t an average pulled from the air. A representative population of historical invoices (every calibrated pair, both sides, a spread of windows, creation dates stepping through the years of price history) is replayed through the same frozen engine the live product runs, and the results are pinned to a fixed artifact (a data file with its own checksum and a stated method). The replay reads only that file.
We show distributions, not averages (the median, the middle 50%, the top 10%) because currency outcomes are skewed and an average hides the tail. Your invoice is placed in an honest band (typical / top-25% / top-10%), never a false “higher than 82%” precision a summary can’t support. One caveat we state plainly: the bank-margin benchmark is model-estimated (fair forward against a typical SA tier), because historical dealt rates aren’t in the set; the measured margin figure is the one you get when you enter your own dealt rate.
They cannot promise the next window. Past outcomes describe how the approaches behaved across thousands of invoices and many kinds of market; any single invoice is close to a coin flip, and the product says so on its own screens. The figures above come from the certification record of the frozen engine: they are not recomputed for marketing, and they will only change when a new sealed test is run and recorded. CNY/ZAR is a synthetic cross (USDZAR × USDCNH); its short-tenor exporter figures are indicative, and execution plans run on USD, EUR and GBP only.