Oisigma/Blog
Research & notes

Overfitting Trading Indicators: Fitted vs Calibrated

Every backtest is a story about the past, and the past is a cooperative witness. Give it enough settings to adjust and it will confirm almost anything. That is overfitting, and it is the single most common reason an indicator that looked exceptional on a historical chart looks ordinary on a live one.

The word gets used loosely, so it is worth being precise about what it is, why it flatters, and what a reader can actually check.

What overfitting is, mechanically

Any indicator has settings: a lookback length, a multiplier, a threshold. Each setting is a degree of freedom, a dial that can be turned. Overfitting, sometimes called curve fitting, is what happens when those dials are turned until the tool agrees with the specific stretch of history in front of it, rather than with the general behaviour of the market.

Any finite stretch of price data contains two things mixed together: durable structure, and noise that happened to fall a particular way in that sample. A tool with enough freedom fits both, and cannot tell them apart, because on the data it was tuned on they look identical. The noise only reveals itself on data the tool has never seen — which is precisely the live market.

So overfitting is not a mistake that carefulness avoids. It is a property of the procedure: if the settings were chosen by looking at the result, the result no longer counts as evidence, however honest the person doing the looking.

The settings search is a hidden parameter

Here is the version that catches people who would never describe themselves as curve fitters.

A trader adds a band to a chart at its default settings. It looks a little wide, so they shorten the window; better. A slightly larger multiplier; better still. After a dozen adjustments the band hugs the recent swings beautifully, and nothing about the process felt like cheating — each step was just making the chart look right.

But "look right" was judged against the same history the settings were chosen on. Every retry was a hidden parameter — a dozen extra degrees of freedom that never appear in the settings panel. The chart now fits for the same reason a key cut to one lock opens that lock. The mechanism is the same whether the search was run by an optimiser or by eye, and by eye is harder to notice, because there is no log of how many things were tried.

This is how confirmation bias operates in backtesting without anyone deciding to be biased. Disagreement gets one more retry; agreement gets a screenshot.

Calibrated is the opposite of fitted

The way out is not fewer settings for their own sake. It is a target stated before the measurement.

A fitted tool is judged by how good the past looks once the settings are chosen. A calibrated tool is judged against a number written down first — a claim the construction makes by design — and then measured on data it was not shaped by. The difference is the direction of the arrow: fitting adjusts the tool toward the data; calibration holds the tool still and asks the data whether the claim was true.

This is why a calibrated result tends to be a modest-looking number. A band built as a one-standard-deviation range makes a specific promise before it ever touches a chart: roughly two-thirds of closes should land inside it. On the S&P 500's daily closes from 1928 to 2024, the measured rate was 71.2%, against a finite-sample benchmark of 67.46% for a 60-bar window. That is a historical measurement, and past behavior is not a guarantee of future results. What makes it evidence is that the target existed first, the construction could not be adjusted to hit it, and the small gap between target and result is itself informative — the paper attributes it to the heavy-tailed shape of real returns, and checked that by running the same construction on simulated markets with and without fat tails. A fitted tool has no such gap to examine; it was tuned until there wasn't one.

What a headline accuracy figure means when it arrives with no stated target is the subject of the holy grail fallacy; this article is about the other half of the question, which is whether the process that produced the number could have manufactured it.

The checks that distinguish the two

A reader cannot see inside a vendor's research process. What they can do is ask which of the following were done and published; each closes a specific route by which a flattering number could have been produced.

A construction fixed in advance. Our own band is a rolling mean and standard deviation of recent returns, projected from the prior close. There is no fitting step, and the one genuine choice — the 60-bar window — is documented with its sensitivity to other lengths in the working paper rather than presented as the winner of a search. What the window length measures is rolling volatility's subject, not this one's.

Causal scoring. Every bar is scored using only information available at the prior close, so the band being tested never contains a glimpse of the close it is tested on. Backtests fail this quietly, through an indicator that recalculates on the current bar or a smoothing step that reaches forward. No lookahead is a prerequisite, not a bonus.

Out-of-sample splits and names the tool never saw. If a construction was quietly shaped by the data it was developed on, it should do worse on data it wasn't. The paper cut each instrument's history in half and compared the two, and separately scored 79 randomly drawn stocks the model was never tuned on; both are set out in the cross-asset evidence, and the paper also reports a separate 2025–2026 holdout period.

A null that re-selects the sample. When a rule chooses which events it counts, a bootstrap confidence interval can look decisive while measuring the wrong thing, because it holds that selection fixed. A rotation null — which re-runs the selection itself — is the stronger check, and it is the reason several results that looked significant in our own testing were reported as nulls.

The misses, published. A calibrated tool's failures are part of its evidence. The Proof page lists where our band runs narrow, where the outer band is optimistic, and what it does not measure at all. Tuning removes the visible misses from the sample it was run on — and only from that sample.

None of these is exotic. Whether a vendor publishes them is something anyone can verify; we do, and the reason is set out here.

What these checks do not establish

A construction surviving all of the above is protected against one specific failure: its coverage claim was not manufactured by tuning. That is a real protection and a narrow one.

It says nothing about whether any use of the tool is profitable. Calibration is a property of a description — how often the next close historically landed inside a stated range. A rule built on top of that description has its own settings, its own search and its own capacity to overfit, and none of the checks above were run on it. The working paper validates the range's calibration, not the profitability of any particular way of using it; whether any such use delivers value after real-world costs is an open question.

Nor does it make calibration perfect in every moment. The band is well-calibrated on average, under- and over-contains on individual stretches, and runs narrow at the onset of a fast volatility spike, in varying degrees, as any backward-looking window does. Those are published limits, not fitting errors.

The takeaway

Overfitting is what any tool with adjustable settings does when the settings are chosen by looking at the result. The defence is not fewer settings but a different order of operations: a claim stated first, a construction that cannot be adjusted toward the data, scoring that never sees the future, tests on data the tool was not shaped by, and the misses published next to the hits.

The most useful question to ask of any indicator, including ours, is not "how accurate is it?" but "what did it claim before it was measured, and on what data was it measured afterwards?"

Frequently asked questions

Is optimising indicator settings the same as overfitting? It can be, and the deciding factor is what the optimisation was judged against. Choosing a setting because it produced the best-looking result on the history being examined is the definition of fitting to that history, whether the search was automated or done by eye. Choosing a setting for a reason stated in advance — a window that spans roughly a quarter, say — and then documenting how results vary across other lengths is a different act, because the choice was not steered by the outcome.

Can a simple indicator with almost no settings still be overfit? Yes, in two ways. The few settings it has can still be tuned to a sample, and the person using it can add hidden degrees of freedom by choosing which chart, which period and which timeframe to show. Simplicity limits the amount of fitting a tool can do; it does not remove the need to test it on data it was not shaped by.

Was the 60-bar window in your own band chosen by backtesting? It was chosen for a stated reason rather than as the peak of a search: 60 daily bars is roughly a calendar quarter, long enough to keep the estimate stable and short enough to adapt to a regime change within about a quarter. The sensitivity of containment to other window lengths is documented in the working paper's appendices, and the figures are historical; past behavior is not a guarantee of future results. The rationale is set out in the FAQ.

If an indicator passes these checks, does a strategy built on it avoid overfitting too? No. The checks establish that the indicator's own coverage claim was not manufactured; a strategy layered on top introduces new settings and a new search, and has to be tested on its own terms. Calibration of a description and profitability of a rule are separate questions, and only the first has been measured in our published work.

The most direct way to judge whether a range was calibrated rather than fitted is to watch it on charts it was never tuned for. BTM draws a calibrated expected range on any TradingView chart, and you can start a free 30-day trial to check it against the markets you already follow.

Now, your charts

Curious how this looks on your charts?

Try it free for 30 days and see the range update as new bars print, on whatever symbols and timeframes you actually trade.

Start your free trial
Complete checkout
Read the paper

30 days free, then $15/mo. Cancel anytime from your account.

Pick up where you left off.