Methodology

Out-of-Sample Testing

It is the single most important idea in all of trading research, and one of the easiest to explain badly. In one sentence: in-sample is the data you used to find a pattern; out-of-sample is data the pattern never saw — and only the second one can tell you whether you found something real. This page explains why that distinction decides whether a statistic is worth anything, and exactly how every HIE report puts it to work.

Key takeaway

In-sample is the history used to find a pattern; out-of-sample is history the pattern never touched. A pattern that only fits its in-sample data has proven nothing — the data can't fail a test it was used to write. HIE splits every report at a fixed boundary (pre-2024 in-sample, 2024-onward out-of-sample) and reports whether the pattern held, weakened, or changed on the unseen data. That check is the difference between evidence and a curve-fit coincidence.

Published
Jul 1, 2026
Last reviewed
Jul 1, 2026
Research through
July 2026
Reading time
7 min
Difficulty
intermediate
Markets
General
Author
Dhaval Barot, MPM Markets
Publisher
MPM Markets
Version
v1.0

It is the single most important idea in all of trading research, and one of the easiest to explain badly. In one sentence: in-sample is the data you used to find a pattern; out-of-sample is data the pattern never saw — and only the second one can tell you whether you found something real. This page explains why that distinction decides whether a statistic is worth anything, and exactly how every HIE report puts it to work.

In 30 Seconds

  • In-sample = the data a pattern was found in; out-of-sample = data it never saw.
  • Only out-of-sample tests anything — a pattern can't fail a test written from the data it was found in.
  • HIE splits every report at a fixed boundary — pre-2024 in-sample, 2024-onward out-of-sample.
  • Three verdicts: held, weakened, changed — telling you how much to trust the finding.
  • This is the defence against "83% win rate" screenshots — those are almost always in-sample only.

Definition

Take any stretch of market history and search it for patterns, and you will always find some. That isn't skill — it's arithmetic. Search enough combinations of conditions and some of them will appear to fit the past purely by chance, the way some lottery ticket always wins. A pattern that fits the data it was discovered in proves nothing, because the data cannot fail a test that was written from it.

So the honest question is never "does this work on the history I searched?" It is "does it keep working on history it never touched?"

  • In-sample data is the history a pattern was found or fitted on. Performance here is almost meaningless on its own.
  • Out-of-sample data is history the pattern never saw during its discovery. Performance here is the real test — the only one that can distinguish a genuine pattern from a coincidence dressed up as one.
In-sampleOut-of-sample
Found the patternTested the pattern
Familiar historyUnseen history
Can always look goodCan genuinely fail
Describes fitTests robustness

Why It Matters

Think of studying for an exam using past papers. If you memorise the answers to papers you've already seen and score 100% on those exact papers, you've proven nothing about whether you understand the subject — you've proven you can memorise. The real test is a new paper you've never seen. Score well there, and you've shown you actually learned something.

A trading pattern is identical. A strategy tuned until it looks brilliant on the years used to build it has memorised those years. Whether it learned anything only shows up on years it was never allowed to see. This is why out-of-sample testing is the single most important idea in trading research: it is the one check that separates a real edge from an expensive illusion.

The practical payoff is blunt. Almost every "83% win rate!" backtest screenshot on the internet is in-sample only — the numbers were computed on the very data that was searched to find the strategy — which is exactly why they tend to collapse the moment real money trades them. A trader who learns to ask "was that out-of-sample?" about every statistic they ever meet — in a research report, in a chat group, in a strategy for sale — has acquired one of the strongest defences there is against buying curve-fit garbage.

How HIE Implements It

Every HIE report performs this check automatically, and the discipline is in the details:

  • A fixed holdout boundary. History is split at a set date: pre-2024 is in-sample, 2024-onward is out-of-sample. The pattern's behaviour is measured on each side separately.
  • The boundary is the same for every query, and it doesn't move. Nobody — not you, not the engine — gets to slide the split until something passes. A holdout you can move until you like the result isn't a holdout. The fixed boundary is what makes the check mean something.
  • The check is computed fresh in every report. It isn't a stored grade; each report re-runs the in-sample-versus-out-of-sample comparison on the specific pattern you asked about.

A real example, from a live report: the ES quiet-volume study split into 394 cases before 2024 (up 52% versus a 54% base rate) and 334 cases from 2024 on (up 53% versus 54%). The behaviour on the unseen data matched the behaviour on the seen data — the verdict was held. That agreement across the boundary is a genuine point in the finding's favour.

(Figures generated 16 July 2026, research through that date; counts can change after a data refresh.)

This is the same principle behind MPM's venue-level Consistency Score badge, applied at a different level: the badge asks whether a whole market holds out-of-sample; this per-report check asks whether this specific pattern held. Same weapon, two targets.

The Three Verdicts

HIE reports the out-of-sample result as one of three verdicts, and each carries a different instruction:

  • Held — the pattern repeated on data it had never seen. This is evidence (not a guarantee) that you found real behaviour rather than a coincidence. Weight the report's numbers as measured.
  • Weakened — usually the recent sample thinned out: fewer recent cases, not necessarily a broken pattern. The honest translation is that less recent confirmation exists, so hold the finding more loosely.
  • Changed — the dangerous one, and the one HIE flags most carefully. A real example: the NQ "RSI below 30" study leaned one direction pre-2024 and then flipped on 2024-onward data — while the typical move size stayed similar. HIE's wording is deliberately precise in cases like this: weight the direction lightly, not the whole path. A directional tilt that doesn't repeat out-of-sample is the classic signature of overfitting — a pattern that fit one past era by chance.

The year-by-year breakdown in a report is the same honesty at finer resolution: it exposes one-year wonders that a single pooled average would quietly bury. If a lean flips from year to year, there is no steady advantage, whatever the overall number says.

Common Mistakes

"It backtested at 80% — that's a great strategy."
Only if that 80% was out-of-sample. An in-sample-only result is computed on the data used to find the strategy, and tells you almost nothing about the future. The first question about any backtest is: was it tested on data it never saw?
"The pattern held out-of-sample, so it's guaranteed to keep working."
No. Out-of-sample survival is evidence, not immunity. Markets can genuinely change after any test; "held through 2024–2026" doesn't bind 2027. The check removes patterns that already failed — it can't see the future.
"'Changed' means the whole report is worthless."
Not necessarily. Often the move size still held while only the direction flipped. A "changed" verdict is a precise instruction to weight the direction lightly — not a reason to discard the path statistics.
"I can just pick the split date that makes my strategy look best."
That's the opposite of a holdout. The moment you move the boundary until something passes, the test is meaningless. A valid out-of-sample check uses a fixed boundary chosen in advance — which is exactly why HIE's is the same for every query and doesn't move.

Limitations

Out-of-sample survival is evidence, not a guarantee. It tells you a pattern held on the specific unseen data it was tested against; it cannot promise the pattern will survive conditions that haven't happened yet. This is why HIE keeps directional readings as context even on findings that held — the check filters out patterns that already failed, but the future is not in any sample.

The check is also only as strong as the data on each side of the boundary. A thin out-of-sample window tests a pattern against few recent cases, which is why "weakened" often reflects a shrinking recent sample rather than a broken pattern. As with every statistic, sample size on both sides of the split governs how much the verdict can bear.

A Habit Worth Keeping

The most valuable thing on this page isn't about HIE at all. It's a question to carry into every piece of trading research you ever read, from anyone:

If the answer is no — or the author never says — you've just found the biggest limitation of that result, whatever its headline number. This one question will sharpen how you evaluate trading systems, research papers, indicators, videos, and strategy advertisements long after the details of this page have faded. It's the single most portable habit in all of trading research.

How This Fits Into MPM

Out-of-sample testing is the backbone of MPM's entire research posture: the belief that a statistic is only worth showing if it has survived data it was never fitted on.

You'll encounter it in:

  • HIE (Historical Intelligence Engine), where every report runs a fresh in-sample-versus-out-of-sample check on your pattern and reports the verdict — held, weakened, or changed — so you know whether the finding survived unseen data.
  • The Consistency Score badge, which applies the same out-of-sample principle at the level of a whole market and timeframe rather than a single pattern.
  • Research Papers, where MPM publishes whether findings held out-of-sample — including the ones that didn't, because a null result that's honestly reported is worth more than a positive one that isn't.
  • Intelligence Circle — a member-only research environment containing advanced market research, historical investigations, and trading strategies developed using the MPM research framework.

MPM's broader stance is captured in one habit: never ask "does this fit the history I searched?" Ask "does it hold on history it never saw?" Out-of-sample testing is that question, made mechanical.

Frequently asked questions

It's data a pattern was never allowed to see while it was being found. Testing on it is like sitting a new exam paper instead of re-marking the ones you memorised — it's the only test that shows whether you actually learned something.

Because a backtest computed on the same data used to find the strategy will almost always look good — that's arithmetic, not edge. The number only means something if it was measured on data the strategy never saw.

It splits history at a fixed boundary — pre-2024 in-sample, 2024-onward out-of-sample — and compares the pattern's behaviour on each side, fresh in every report. It then reports whether the pattern held, weakened, or changed.

Held: the pattern repeated on unseen data — real evidence. Weakened: usually a thinner recent sample — hold it loosely. Changed: the pattern (often its direction) flipped out-of-sample — the signature of overfitting; weight it lightly.

Because every HIE report must apply the same transparent test. A fixed holdout is simple to state, impossible to game, and identical for every query — which is exactly what makes the result trustworthy. The point isn't a clever validation scheme; it's a test nobody can bend.

The important property isn't the specific year — it's that the boundary is fixed in advance and the same for every query. A holdout chosen before questions are run can't be tuned to flatter a result; that fixed, unmovable quality is what makes it a real test.

No. It's evidence the pattern survived the unseen data it was tested on — not a promise about the future. Out-of-sample survival reduces uncertainty; it doesn't eliminate it. Markets can still change.

Same idea, two levels. The Consistency Score asks whether a whole market and timeframe holds out-of-sample in general; this per-report check asks whether your specific pattern held. One grades the ground, the other grades the idea standing on it.

Supporting evidence

Research, methodology and datasets supporting this page.

Where you'll encounter this

Continue your research journey

Methodology
Consistency Score

Before you trust any statistic, there's a question worth asking that almost no tool asks: should this number exist at all? MPM's Consistency Score is the badge on every HIE report that answers it. It doesn't grade the pattern you asked about — it grades the ground you're standing on: whether statistics measured on this market and timeframe have a track record of holding up on data they were never measured on. This page explains what the badge means, why it exists, and the important two-layer distinction behind it.

Trading Statistics
Sample Size in Trading

Sample size is simply how many trades — or how many observations — a statistic is based on. It's the least glamorous number in trading and arguably the most important, because every other metric is only as trustworthy as the sample behind it. This page explains what sample size is, why a small sample can make almost any result look good, the mistakes people make with it, and why sample size sits at the centre of how MPM decides whether a number counts as evidence.

Statistics
Z-Score

Today's move is bigger than usual — but how much bigger? A z-score answers exactly that, by measuring a value against its own recent history and expressing the gap in standard-deviation units. A z-score of 0 means "exactly average"; +2 means "two standard deviations above average — an unusually high reading"; −2 means "unusually low." It's one of the cleanest ways to turn a raw number into a statement about how rare it is. This page explains what a z-score is, why it's useful, how HIE computes it, and the mistakes people make reading it.

Statistics
Percentile

Is today's trading volume high? A percentile answers that in plain terms: it tells you what share of recent readings were lower than today's. If volume is in the 95th percentile, today was higher than 95% of recent sessions — genuinely elevated. If it's in the 40th percentile, it was fairly ordinary. This page explains what a percentile is, why it's often more honest than an average, how HIE computes it, and the mistakes people make reading it.

Citations

  1. MPM Markets (2026). Out-of-Sample Testing. MPM Learning Center.Suggested citation: MPM Markets (2026). Out-of-Sample Testing. MPM Learning Center. mpmmarkets.com/glossary/out-of-sample-testing

Suggested citation

Dhaval Barot, MPM Markets (2026). Out-of-Sample Testing. MPM Markets Retrieved from https://mpmmarkets.com/glossary/out-of-sample-testing

Reviewed Jul 1, 2026 · Research current through July 2026 · v1.0