Oisigma/Blog
Research & notes

Do Order Blocks and Liquidity Sweeps Actually Work?

Smart-money concepts is one of the most widely taught frameworks in retail technical analysis, and one of the least measured. Its building blocks — order blocks, liquidity sweeps, fair-value gaps — are usually presented as descriptions of what large institutions are doing. That is a claim about causes, and causes are close to untestable from price data alone.

So we did not test the causal story. We tested the outcome claim each construction implies, which is narrower and answerable: does this pattern mark bars or levels whose forward behaviour differs measurably from a matched comparison?

Two of the three primitives have now been run, on 37 instruments of daily data. Both came back null on selection. One of them also produced a calibration result that went against our own model, which is reported here for the same reason we report the rest.

Read this limit first

These runs used daily bars, and smart-money concepts is largely practised intraday. That is a real limitation rather than a footnote, and it bounds everything below. A trader working five-minute charts can fairly say these results were not measured on their timeframe, and they would be right.

What follows establishes that the daily-bar versions carried no selection information in this sample. Whether the intraday versions do is untested here. One more scope note: order blocks and liquidity sweeps were run; the fair-value-gap primitive was specified and not run.

Order blocks: the gate carried the information, the pattern did not

An order block is the zone left behind by a large "displacement" bar. The implied claim is that price returns to that zone and behaves differently there than at an arbitrary level — revisiting more often, or producing a more favourable excursion in one direction.

Two questions were registered in advance. First, whether gating order blocks on a calibrated displacement threshold changed their base rates against an ATR threshold at matched event rates. Second, and more interesting: whether the order-block geometry added anything at all over the displacement bar it is drawn around. That comparison ran 4,108 gated order blocks against 31,435 ungated candidates.

Gating on magnitude did change the outcome distribution. The full construction beat ungated pattern geometry by +0.361 [0.074, 0.611], with the interval excluding zero at every threshold tested.

The pattern itself added nothing. Measured against the bare gated bar, the difference was −0.276 [−0.813, 0.161] — the kill condition stated before the run, and met. At the most selective threshold, where the comparison is sharpest, the interval excluded zero on the negative side: −0.679 [−1.363, −0.096]. Where the geometry made a measurable difference, it made the outcome worse.

On base rates the answer was flat. Against a null built from 400 circular rotations of the gate, the critical statistic was 8.70; the largest value observed anywhere on the 60-cell grid was 6.00, the 77th percentile of that null. Nothing cleared.

The honest summary: the object carrying information was the abnormally large bar. The zone drawn around it was decoration.

Liquidity sweeps: no selection, but a bigger move

A sweep is a bar that pierces a prior extreme and then closes back inside it — in the usual telling, a stop run that precedes a reversal. The testable version is that sweeps should be followed by reversal, and deeper sweeps should reverse harder.

This was the best theoretical fit of the three primitives. Excursion depth past a prior extreme is the containment question in different clothing, which made it worth running properly: 37 assets, 22,169 sweeps, 207,790 bars.

Measuring excess forward return over each asset's own drift, signed toward the reversal, every arm's interval contained zero and every point estimate was negative. Sweeps generally came in at −13.4bp [−44.5, +8.1]; sorting by depth did not rescue it, at −17.4bp for the volatility-normalised arm and −28.8bp for raw percentage depth.

The best-performing arm was the placebo. Shallow sweeps — the ones the hypothesis says should be least reversal-prone — landed at −3.5bp [−29.3, +14.2], and were positive on 24 of 37 assets against 17 of 37 for the real arm. Under the rotation null, nothing survived.

The control endpoint did show something. Forward volatility after a deep sweep ran 1.105× baseline [1.074, 1.140], against 0.973× for sweeps generally. Deep sweeps were followed by a bigger move — not a directional one. That is the same magnitude-not-direction conclusion the rest of our research keeps reaching, and the control was registered specifically to separate the two.

The result that went against our own model

On one endpoint, our own volatility measure lost.

Comparing how consistently a single global threshold produced the same event rate across instruments — the calibration property we normally argue for — ATR-normalised depth beat standard-deviation-normalised depth at every comparable setting. The cross-asset coefficient of variation was 0.324 for ATR against 0.564 for ours at matched rates, and the gap held at tighter settings.

The mechanism resolves it rather than explaining it away. Excursion depth past a prior extreme is an intraday quantity: it is measured from a bar's high or low, not its close. ATR's denominator contains the intraday range, so it is dimensionally the right normaliser for it. A measure built from closing prices is not.

The rule the two results jointly establish is sharper than either alone: standard deviation calibrates close-to-close quantities; ATR calibrates intraday ones. That is a real boundary on how far our own calibration property extends, and it means any claim of a general calibration advantage over ATR is wrong as stated. We would rather publish the boundary than leave the broader claim standing — which is the same reason we publish the crisis-onset limit.

Why the bootstrap kept saying yes

One pattern recurred often enough across these runs to be worth stating on its own.

Three of five bootstrap confidence intervals in the order-block run excluded zero. The rotation null killed all three. It happened again in the sweep run: one comparison showed +11.4bp with a bootstrap interval of [+0.5, +24.5] that excluded zero, and a rotation p of 0.65.

The reason is structural. A cluster bootstrap resamples assets and time while holding the event selection fixed. A rotation null re-selects which events the rule admits in the first place. For any quantity where a rule chooses its own sample, the bootstrap standard error is the wrong yardstick — here it understated the true variability by roughly a factor of three.

This is a finding about method rather than markets, and it travels beyond our own work: any result resting on a bootstrap interval over a rule-selected quantity, with no rotation null behind it, is weaker than it looks.

What this does not establish

It does not establish that these constructions fail intraday, which is where they are mostly used. It does not cover fair-value gaps. It does not establish that the calibration advantage found elsewhere in our work is unique to our model. And it makes no economic claim in either direction — no transaction costs were modelled, because nothing here converts into a performance statement without a further run.

It says nothing about any particular commercial indicator. What was tested is the public-domain definition of each construction, implemented from its standard description.

Most importantly, it does not establish that traders using these patterns are wrong to find them useful. A pattern can organise attention, enforce consistency, or encode a discretionary skill its written definition does not capture. What these runs measured is narrower: whether the construction, applied mechanically on daily bars, selected a distinguishable forward distribution. In this sample, it did not.

The takeaway

Both tests pointed the same way. What carried information was the size of the move relative to what was normal for that market. What carried none was the geometry drawn around it afterwards.

That is not an argument that our own band is the answer instead. It is an argument for asking any tool the same question — what does this claim, and has anyone measured it? — and for publishing the answer when it comes back inconvenient. One of the two results here is inconvenient for us.

Frequently asked questions

Does this mean order blocks don't work on 5-minute charts? No. These runs used daily bars only, and that limit is stated up front precisely because most smart-money-concepts trading happens intraday. Nothing measured here transfers to a five-minute chart in either direction. Establishing whether the intraday versions carry selection information would require a separate run on intraday data, which has not been done.

What about fair value gaps — were those tested? Not yet. Three primitives were specified at the outset; order blocks and liquidity sweeps were run, and the fair-value-gap primitive was defined but never executed. It is listed here rather than quietly dropped because a reader is entitled to know which parts of a stated plan were completed.

Does this prove smart money concepts is fake? It does not, and the claim would be much wider than the evidence. What was tested is whether two specific constructions, applied mechanically on daily bars, selected forward distributions that differed from a matched comparison. They did not, in this sample. That is compatible with the framework being useful to a discretionary trader for reasons the mechanical definition does not capture, and it says nothing about the primitive that was not run.

Is ATR better than standard deviation, then? For intraday quantities in this test, yes — and that is worth stating plainly rather than hedging. The finding is scoped: ATR calibrated excursion depth past a prior extreme better than a close-based measure at every comparable setting, because depth is measured from highs and lows and ATR's construction contains that range. For close-to-close quantities the ranking runs the other way. Neither result generalises past the type of quantity being measured.

If you want to see what a construction looks like when the claim it makes is stated in advance and then checked, the working paper documents the method, the figures and the places the model is weakest, and the Proof page summarises them. The indicator itself is a 30-day free trial if you would rather watch it recalculate on your own charts than read about it — you can start one here.

Now, your charts

Curious how this looks on your charts?

Try it free for 30 days and see the range update as new bars print, on whatever symbols and timeframes you actually trade.

Start your free trial
Complete checkout
Read the paper

30 days free, then $15/mo. Cancel anytime from your account.

Pick up where you left off.