Ask a trader whether gold behaves like Bitcoin, or whether the S&P 500 behaves like a ten-year Treasury future, and the answer is obvious. They do not. They move by different amounts, on different schedules, driven by different participants; a typical day in one looks nothing like a typical day in another.
That answer is correct, and it answers a question about magnitude. A second question hides underneath it, and the two get conflated constantly: not how far does this market travel, but how well does a range built from this market's own recent behaviour describe where its next close lands. Those are separate properties, and a market can be wild on the first while being entirely ordinary on the second. Our working paper put the second question to forty instruments, and the result is the more interesting kind — not a dramatic number, but a stubbornly consistent one.
The intuition that markets each need their own settings is reasonable and widely held: volatility differs by orders of magnitude across asset classes, so a parameter that suits a currency pair intuitively ought to be wrong for a small-cap stock.
But that intuition is about the scale of the numbers, and scale is a property of the unit you measure in, not of the market. A band expressed in pips, points or dollars carries the price level with it. A band expressed in percentage returns does not — it is rebuilt on every bar from the dispersion of that instrument's own recent percentage moves, then projected back onto price from the prior close. Two instruments quoted at completely incomparable prices have daily percentage moves that are comparable. That measurement choice is the load-bearing one, and it is worked through in return space vs price space. Once every instrument is measured in its own returns, one question can be asked of all of them at once.
The question is containment: given a band drawn from an instrument's recent volatility, did the next close land inside it, or outside? It is answered bar by bar, and it produces a percentage rather than an opinion. What "calibrated" means here is set out in our pillar post, and the mechanics of the test in how to compare band indicators.
Two properties of the scoring matter for a cross-asset claim. It is strictly causal: for each bar, the band is built from information available at the prior close only, and the realised close is then classified as inside or outside. And nothing was fitted per instrument — the same rolling calculation, with the same 60-bar window, was run on gold, on a Treasury future and on a single stock, with no per-market tuning to make any of them agree.
The tested universe is 40 instruments spanning five asset classes — equities including single names and sector ETFs, FX, commodities, rates, and crypto — measured on daily closes over each instrument's own available history, totalling roughly 213,000 close observations.
Across that universe, the average share of closes landing inside the inner band was about 71.65%, and inside the outer band about 94.0%. The dispersion is the part worth pausing on: the standard deviation across instruments was roughly 2.06 percentage points at the inner band and 0.7 points at the outer one, with most instruments falling between about 68% and 77%. All of these figures are historical and were measured in research; past behavior is not a guarantee of future results. Every published row, with its sample period attached, sits on the Proof page.
Forty markets that share almost no other property landed within a few points of each other on this one. The asset-class detail behind that is covered in Bitcoin's expected range vs stocks and forex expected ranges rather than repeated here.
A sharper version of the consistency check sits inside the table, and it is the one we find most persuasive.
The same underlying exposure trades in several different forms: the cash index, which is not directly tradable; the ETF built on it; the futures contract. Three separate data series, with different histories, trading hours, microstructure and holders — and economically, the same market underneath. Scored on containment, they calibrated alike. That is a harder result to explain away as an artifact of one dataset, because the three series were built and cleaned independently of one another and still produced the same answer.
The obvious objection to any containment number is circularity. If the band is drawn from the instrument's own volatility, of course price stays inside it — the result is baked in.
The working paper tested exactly that by running the identical method on thousands of simulated markets. On a plain, well-behaved Gaussian market, the same construction with the same finite window produced containment of roughly 67%. That is the benchmark the circularity objection predicts. On simulations built to reproduce the fat-tailed behaviour of real returns, it produced roughly 71% — matching the live cross-asset result to within about a tenth of a percentage point. These figures are historical and were measured in research; past behavior is not a guarantee of future results.
So the objection has a real mechanism behind it, but the extra few points are not free. They are the fingerprint of how actual return distributions behave.
Consistency across instruments still leaves open the possibility that the construction was quietly shaped by the instruments it was developed on. Two checks address that.
The first is an out-of-sample split: each instrument's history was cut in half, and containment in the second half compared with the first. It was essentially unchanged.
The second is a fresh set of names. On 79 randomly drawn U.S. stocks the model was never tuned on, containment averaged about 73.6%, and every single name cleared 70%. Both figures are historical; past behavior is not a guarantee of future results. The random draw was, if anything, the easier test — the diverse, curated universe turned out to be the harder one.
Consistent average calibration is a narrower claim than it can be made to sound, and several published limits sit next to it.
It is calibration on average, not in every moment. The band is a marginal-coverage description, not a perfect conditional forecast, and it under- and over-contains on individual stretches. The outer band runs slightly optimistic in the deep tails. And a band built from a rolling window of recent history runs narrow in the opening days of a fast volatility spike, in varying degrees, because the spike has not entered the window yet — in our own testing, on the most extreme crisis-onset days, inner-band containment fell to roughly 65% against about 71% in ordinary conditions; both figures are historical and past behavior is not a guarantee of future results. That mechanism is set out in when volatility bands fail.
It also says nothing about direction, and nothing about profit. The paper validates the range's calibration, not the profitability of any particular way of using it; whether any such use delivers value after real-world costs is an open question, and trading always carries the risk of loss.
Markets differ enormously in how far they travel. That is the question most volatility tools answer, and pips, points and dollars are reasonable units for it. Scored instead on whether a range built from each market's own recent behaviour described where its next close landed, the tested universe — plus three separate wrappers of the same exposure, an out-of-sample split, and 79 names the model had never seen — came back with answers a few points apart. The consistency, not the exact percentage, is the result. Past behavior is not a guarantee of future results.
Does it work on markets that weren't in the 40 tested? The construction makes no assumption about which market it is running on, so it computes anywhere a series of closes exists. The evidence, though, covers what was measured. Instruments outside the tested universe are untested rather than validated, and the reasonable expectation is that behaviour transfers best where returns are reasonably continuous and liquidity is decent.
Which market does it work best on? The working paper does not publish a ranking, and treating the spread as one would be reading more into it than it supports. The instrument-level differences fell inside a spread of roughly two percentage points at the inner band in the tested sample, and past behavior is not a guarantee of future results. A ranking would also not mean what rankings usually imply — calibration is a description of fit, not a measure of performance.
Do the same settings really apply to every market? The published cross-asset figures were produced with one configuration — a 60-bar volatility lookback, roughly a calendar quarter on daily bars — applied identically to every instrument, with no per-market fitting. Shorter windows react faster but noisier; longer ones smooth more but lag through regime changes. Window sensitivity is documented in the working paper rather than left to preference.
Is the working paper peer-reviewed? No, and we label it a working paper deliberately so readers know what they are reading. It is complete and citable, the construction is fully specified, and the results are reproducible from the documented data and methods — but it has not been refereed by an academic journal. Anyone is welcome to check it.
The full cross-asset table, the sample period behind every row, the simulation benchmarks and the published limits are all on the Proof page, with the derivations in the working paper — and the most direct way to judge a cross-asset claim is to watch the same construction recalculate on markets that have nothing in common. You can start a free 30-day trial and put it on a stock, a currency pair and a commodity chart at the same time.
Oisigma provides descriptive market analytics for educational use. It is not investment advice, does not predict prices, and does not provide buy or sell signals. Statistics referenced are historical and were measured in our working paper (not peer-reviewed); past behavior is not a guarantee of future results. Trading and investing involve substantial risk of loss, including the possible loss of all capital invested. Leveraged products (futures, options, margin) carry additional risk and can result in losses that exceed your initial investment. Bollinger Bands® is a registered trademark of John Bollinger; Oisigma is not affiliated with or endorsed by Mr. Bollinger. RiskMetrics® is a registered trademark of MSCI Inc.; Oisigma is not affiliated with or endorsed by MSCI Inc. Nothing in this article is a recommendation to use any particular strategy. Read the full Disclaimer →
Try it free for 30 days and see the range update as new bars print, on whatever symbols and timeframes you actually trade.
30 days free, then $15/mo. Cancel anytime from your account.
Pick up where you left off.