Oisigma/Blog
Research & notes

The Holy Grail Fallacy in Trading Indicators

Search for the most accurate trading indicator and the results arrive with numbers stapled to them: 95% accuracy, 99% win rate, catches the exact top and the exact bottom. Scroll a little further and the same search surfaces the other half of the internet — forum threads and comment sections where traders who have already been down that road answer with near-total unanimity that the holy grail does not exist.

Both halves are missing the same thing. The sellers publish a number without saying what it counts. The skeptics conclude that numbers attached to indicators are worthless in general. The more useful position sits between them: some numbers attached to an indicator do mean something specific and checkable — and the way to tell which is not how large the number is, but whether it was aimed at a target stated in advance.

Accuracy of what, exactly?

“95% accurate” reports a rate without naming the event being counted. That omission is the whole trick, because at least three very different numbers travel under the same word.

A win rate counts how often trades taken from a tool’s signals ended in profit. It depends on the entry rule, the exit rule, the holding period, the costs, and the slippage — none of which are properties of the indicator. Change the exit and the number changes. Two people running the same script honestly report different figures.

A hit rate on a labelled event counts how often something the tool flagged was followed by something the tool defines as success. It depends entirely on where the labels were drawn, and the labels are usually chosen after looking at the data.

A containment rate counts how often the next close landed inside a range the tool drew before that bar existed. There is no entry, no exit, no cost assumption, and nothing to tune after the fact. It is a property of the drawing itself.

Only the third kind survives being handed to a stranger. The first two are joint claims about a tool and a trading process, presented as claims about a tool.

Calibration is the unglamorous word for what matters

A containment rate becomes interesting the moment you compare it to what the band claimed. That comparison is called calibration: does the rate a construction implies match the rate it actually delivered? A band that implies 70% and delivered close to 70% is calibrated. A band that implies 95% and delivered 83% is not — and the point isn’t that 83% is a poor number, it’s that the 12-point gap is where an unsuspecting reader gets surprised. The full definition, and why it belongs on the label, is the subject of its own post.

What makes this testable rather than rhetorical is the ordering. The claim is fixed before the bar prints. The outcome is counted after. Nothing in between can be adjusted.

Why a matched number is harder to produce than a large one

Here is the part the holy grail framing gets backwards. In a band indicator, largeness is free. Width is a dial. Widen a band far enough and it will contain 100% of next closes while describing nothing at all — a range from zero to infinity is never wrong and never useful. Anyone can manufacture a high containment number in an afternoon.

What cannot be manufactured is a number that lands where the construction said it would.

Across 97 years of S&P 500 daily closes, the inner band described in Oisigma’s working paper contained the next close 71.2% of the time, against a correct finite-sample benchmark of 67.46% for its construction. That is a small excess, in the direction of slightly more containment than implied, and the paper spends its space explaining the excess rather than rounding it away: it is consistent with heavy-tailed returns, and a heavy-tailed model reproduces it where a Gaussian one does not. The outer band’s cross-asset mean sat near 94%. Past behavior is not a guarantee of future results.

For contrast, scored by that same advance-commitment question on the same data, a standard price-space Bollinger Bands® construction at its canonical (20, 2) settings — which implies roughly 95% containment under a normal assumption — contained about 83%. That comparison is specific to those two constructions at that width, and it is audited at length in its own post, settings and all.

So the smaller headline number is the stronger claim. Not because 71 is better than 95 in any race, but because 71 was aimed somewhere and landed near it, with the residual explained.

The second tell: does the number arrive with its failure conditions?

A calibration figure quoted alone is an average wearing the costume of a guarantee. Averages hide their bad days.

The published record for this band includes its bad days by design. Containment fell to 65.29% when the VIX was above 30 — a structural consequence of building a range from recent behavior, which cannot widen until the new volatility has entered the window. The outer band also runs slightly optimistic in the deepest tails. Neither is a footnote discovered by a critic; both are in the paper and on the evidence page, and the crisis-onset limit has a post of its own.

This is the cheapest test a reader can apply to any tool, including this one. Every construction built from historical data has conditions under which it performs worse, so a number published without any is a number whose limits have either not been looked for or not been reported. The limits exist either way.

What a calibration number still cannot do

Being well calibrated is a narrow virtue, and it is worth being precise about how narrow.

It says nothing about direction. A range describes the likely size of the next move, not its sign; the center line in this construction has historically carried almost no directional information, and is documented as a reference rather than a forecast.

It says nothing about profit. The paper validates the calibration of the range, not the profitability of any use of it. Whether any particular way of using a range delivers value after costs is an open question that a containment study does not address, and describing how some traders use a range is not the same as evidence that doing so works.

And it is an average property, not a promise about any single bar. That distinction is exactly the one that gets lost when a statistic travels — the reason a heavily qualified research finding turns into a flat internet fact, as happened to the most-repeated number in trading.

The takeaway

The holy grail fallacy is not the belief that a tool can help. It is the belief that a bigger advertised number is better evidence than a smaller one. In a band indicator the opposite is closer to true: size is a dial, and the only figure that costs anything to produce is one that had to match a target fixed in advance — published with the sample it came from, the benchmark it was scored against, and the conditions where it falls short.

That reframes the question a reader should put to any indicator, including this one. Not how accurate is it, but: what did it commit to before the data arrived, what did it deliver, and where does it break?

Frequently asked questions

Does the Behavioral Transform Model repaint? Once a bar closes, its range and its markers are fixed and do not change afterwards. While the current bar is still forming, the live abnormal-move marker on that unfinished bar can update until it closes. The calculation uses only data available at the prior bar, which is what makes a containment record meaningful — a construction that could revise its own history could not be scored against an advance commitment at all.

What is a realistic accuracy figure for an indicator? The question is hard to answer as posed, because “accuracy” is not a single quantity — a win rate, a labelled hit rate, and a containment rate measure different things and are not comparable. For a containment rate specifically, realism is judged by the gap between what a construction implies and what it delivered, not by the size of the number. A band can be widened until it contains almost everything and still describe nothing.

Does Oisigma publish a win rate for its indicator? No. The Behavioral Transform Model is descriptive and produces no entries, exits, or signals, so there are no trades to score. What is published is a containment record — how often the next close landed inside the range, measured across markets and periods — together with the regimes where that record weakened. Past behavior is not a guarantee of future results.

How much of the market does the published evidence actually cover? The working paper covers forty instruments across five asset classes — equities, FX, commodities, rates, and crypto — with out-of-sample and holdout tests, plus weekly and monthly bars. The 97-year span applies to S&P 500 daily data specifically; the other instruments have shorter histories determined by when reliable data begins. Anything outside that set is untested rather than validated.

Oisigma’s headline figure was aimed at a benchmark before it was measured, and the regimes where it falls short are published next to it. If that is the standard you would rather judge tools by, the fairest way to check it is against the markets you actually watch — the 30-day free trial exists for that, and the paper is free to read either way.

Now, your charts

Curious how this looks on your charts?

Try it free for 30 days and see the range update as new bars print, on whatever symbols and timeframes you actually trade.

Start your free trial
Complete checkout
Read the paper

30 days free, then $15/mo. Cancel anytime from your account.

Pick up where you left off.