Terminology follows behavior: a period stops being out-of-sample the moment its result influences strategy selection, parameter choice, or whether the idea is abandoned.

Three different jobs for historical data

Training or in-sample period

The in-sample period is used to formulate and fit the strategy. Indicator windows, thresholds, filters, and model parameters may be selected here. Performance in this period answers whether the rule can fit the development data—not whether it generalizes.

Validation period

A validation period can compare candidate designs, tune meta-decisions, or choose between parameter regions. Because its outcomes influence selection, it belongs to the development process even if it was initially called “out-of-sample.”

Final test period

The final holdout is evaluated after the strategy, costs, benchmark, and decision criteria are frozen. It should be inspected once. If the result triggers another development cycle, preserve that result as part of the research history and obtain new future data for genuinely independent evidence.

Why chronological order matters

Randomly shuffling daily market observations destroys the time direction a trading rule faces. Adjacent returns, volatility regimes, indicator windows, and labels can overlap. A randomly selected training row may use observations next to—or even inside the calculation window of—a test row.

At minimum, training must precede testing. If observations use forward return labels or overlapping holding periods, a gap or embargo may be needed at the boundary so training labels do not contain returns from the test interval. The correct gap depends on the longest feature window, signal delay, and label or holding horizon.

A single chronological split

A simple design might use the first 60% of history for training, the next 20% for validation, and the final 20% for one test. Percentages are not universal. The periods must contain enough independent market behavior and enough strategy events to make the chosen statistics interpretable.

A ten-year daily history split 6/2/2 may appear substantial, yet a slow strategy could produce only a handful of trades in the final two years. The effective sample is determined by opportunities and dependence, not merely 500 daily rows.

The advantage is clarity: the final test represents one forward chronological block. The disadvantage is sensitivity to that boundary and regime. A strategy trained in a long bull market and tested during a crash may reveal useful fragility, but one outcome cannot describe every possible transition.

Walk-forward evaluation

Walk-forward testing repeats a train-then-test sequence through time:

CycleTraining informationUntouched evaluation
1Years 1–4Year 5
2Years 2–5 or 1–5Year 6
3Years 3–6 or 1–6Year 7

A rolling window keeps training length fixed and forgets older data. It may adapt to changing regimes but estimates parameters from less information. An expanding window retains all earlier observations. It has more data but assumes older relationships remain relevant.

At each cycle, all fitting—including indicator parameters, feature normalization, cost calibration rules, and strategy selection—must use only that cycle’s training window. The selected frozen rule then produces returns for the next evaluation segment. Those out-of-sample segments are concatenated chronologically into one synthetic record.

Aggregate returns before calculating metrics

Do not average the Sharpe ratios or drawdowns of walk-forward folds. Sharpe is nonlinear, and maximum drawdown can begin in one segment and end in another. Link the out-of-sample daily returns in time order, compound one equity curve, and calculate return, volatility, drawdown, and trade statistics from that combined path.

Record the selected parameters per cycle. Stable neighboring values suggest a broad response surface; violent changes between windows may indicate noise, regime sensitivity, or an underidentified model.

Leakage that survives a date split

  • Full-sample preprocessing: scaling, winsorizing, or filling values using statistics from dates after the training boundary.
  • Present-day universe: using today’s constituents throughout historical folds.
  • Revised data: using a value as later corrected rather than as published at the decision time.
  • Global parameter grids: designing the search range after studying final-period outcomes.
  • Cost hindsight: calibrating friction from information unavailable in the simulated training period.
  • Repeated peeking: retaining only strategies that looked good across the reported out-of-sample segments.

Validation is still model selection

If 40 strategies and thousands of parameter combinations compete, choosing the validation winner creates selection bias. The final test protects against that selection only if it remains untouched. The number of trials and correlation between them determine how surprising a winner really is.

Bailey, Borwein, López de Prado, and Zhu show why ordinary holdout techniques can be unreliable in investment backtests and propose combinatorially symmetric cross-validation to estimate the probability of backtest overfitting. Their framework is more involved than a single split, but the practical lesson is immediate: performance must be interpreted in light of the full selection process.

Predefine the walk-forward design

Window length, step size, metric, minimum trades, rebalance schedule, cost scenarios, and failure criteria are themselves parameters. Searching many walk-forward designs and reporting the best repeats the original problem one level higher.

Write these choices down first. A defensible protocol might state: use an expanding four-year minimum training window, refit annually, evaluate the following year, select only among a fixed list of parameter configurations, require a minimum event count, and aggregate every out-of-sample year—including failed ones.

What TickRun currently does

TickRun requests the latest available ten-year saved history and evaluates the chosen strategy across that full range. Its optimizer searches parameters and ranks them on the same range. Therefore:

  • normal and optimized results are historical in-sample descriptions;
  • the maximum-combination and time limits bound the search but do not create out-of-sample evidence;
  • testing another saved ticker checks cross-asset behavior, not future-time performance;
  • signal delay, minimum holding period, and cost stress can reveal fragility but are not holdouts.

A future implementation should accept explicit train, validation, and test boundaries; lock the final period during optimization; preserve every trial; and build a combined walk-forward out-of-sample equity curve. Until then, users should not label an optimized TickRun result “validated.”

What a validation report should contain

  • Exact chronological boundaries and any embargo.
  • Data snapshot and point-in-time limitations.
  • All candidate strategies and parameter grids.
  • Selection metric and tie-breaking rule.
  • Parameters chosen in each cycle.
  • Every out-of-sample segment, including losing segments.
  • One combined out-of-sample equity curve and metrics.
  • Cost assumptions and stress scenarios.
  • Number of times the holdout was viewed or reused.

Research references