The central problem: a backtest describes one historical sample. If you search enough rules, some will fit that sample unusually well by chance—even when they have no durable advantage.
What is overfitting?
Overfitting happens when a strategy is adjusted so closely to historical data that it captures random details rather than a repeatable market relationship. The finished strategy may explain the past beautifully but fail when prices behave even slightly differently.
Imagine drawing a smooth line through a noisy set of points. A simple line will miss some observations but may describe the broad pattern. A highly flexible line can touch every point, including every accidental bump. That second line has a better historical fit, yet it may be worse at describing the next point. Strategy optimization creates the same tension.
Overfitting does not require dishonest research or a complicated machine-learning model. It can happen whenever choices are influenced by the results: selecting a strategy, changing a lookback, moving a threshold, choosing an asset, altering costs, or deciding which time period to show.
How a backtest becomes overfit
A researcher often starts with a reasonable question, runs a test, and sees an unsatisfying result. The natural response is to adjust a parameter and try again. If the new result improves, another adjustment follows. Eventually, the process may produce an impressive equity curve.
The difficulty is that every result revealed information about the same historical sample. Even if the researcher never writes down a formal optimization loop, repeated trial and error is still a search. The final rule has indirectly learned which settings happened to work in that particular history.
Automated optimization makes this process faster and more visible. TickRun’s default ceiling permits up to 5,000 configurations per strategy. Optimizing all 40 strategies can therefore inspect as many as 200,000 configurations when time limits allow. Selecting the highest return from such a large field is not equivalent to testing one independently chosen hypothesis.
Why the winner is often too good
Suppose thousands of strategies have no genuine predictive value but produce different historical outcomes because their trades occur on different days. Some will lose, many will look ordinary, and a few will appear excellent purely by chance. An optimizer reports the excellent ones because finding the maximum is its job.
This is selection bias: the winning result is shown after the competition, while the thousands of unsuccessful alternatives fade into the background. The displayed return does not contain an automatic penalty for how many candidates were tried.
More combinations can be useful for exploring how a rule behaves, but they do not automatically create stronger evidence. A result found after 100,000 attempts generally needs more independent validation than a result produced by a rule specified before looking at the data.
Common warning signs
A narrow parameter peak
A strategy performs very well with a 17-session lookback but poorly with 16 or 18. Unless there is a convincing reason that 17 sessions should be special, the isolated peak may reflect historical noise. Robust ideas tend to survive small parameter changes, even if their exact returns differ.
High return from very few trades
A small number of fortunate trades can dominate a ten-year result. Five winning trades do not provide the same evidence as a behavior repeated across many independent opportunities. A minimum-trades filter helps, but it cannot guarantee that trades are independent or representative.
Results depend on one asset or market phase
A rule may work only for one company, one sustained bull market, or one volatility event. That can still be an interesting observation, but it is weaker evidence of a general trading effect. Different assets and market conditions expose different ways a rule can fail.
Small assumptions reverse the outcome
If modest transaction costs, a one-session signal delay, or a slightly longer holding constraint destroys the result, the apparent advantage may be too fragile for realistic use. Robustness to reasonable assumptions matters more than preserving the maximum historical return.
An unnaturally smooth equity curve
A nearly perfect historical curve deserves investigation, not automatic confidence. Check whether a few observations dominate, whether the strategy trades unrealistically often, whether costs are set to zero, and whether the rule was selected from a very large search.
How to reduce overfitting risk
Start with a hypothesis
Write down why the indicator might capture a repeatable behavior before optimizing it. For example, a trend rule might be based on persistence after large information shocks, while a mean-reversion rule might be based on temporary overreaction. A plausible explanation does not prove an advantage, but it limits arbitrary searching.
Keep the rule simple
Every adjustable threshold, filter, exception, and timing rule adds flexibility. Complexity may be justified, but it also creates more ways to fit noise. Prefer the simplest rule that expresses the hypothesis and treat added parameters as costs that require evidence.
Reserve genuinely unseen data
The strongest basic safeguard is to design the strategy on one period and evaluate it once on a later period that played no role in the design. If you repeatedly inspect the reserved period and continue adjusting, it gradually becomes training data too.
Walk-forward analysis repeats this idea through time: configure or select a rule using past data, evaluate it on the next unseen segment, and then move the window forward. The combined unseen segments provide a more realistic record of the research process.
Look for parameter stability
Do not inspect only the winning configuration. Examine nearby values. A broad region of acceptable outcomes is generally more credible than one spectacular coordinate surrounded by weak results. Stability does not prove future performance, but instability is a useful warning.
Test different assets and regimes
Ask whether the underlying behavior appears in other companies, sectors, or broad markets. Also consider rising, falling, calm, and volatile periods. The goal is not to demand identical returns everywhere; it is to learn whether the proposed mechanism travels beyond the sample that inspired it.
Use realistic friction
Transaction costs should not default to zero in serious analysis. Real execution can also include spreads, slippage, taxes, market impact, and unavailable closing prices. TickRun models a configurable cost on position changes, but it does not model every source of friction.
Record the full search
Keep track of how many strategies and configurations were tested, including failures. A 40% winner found on the first preplanned test is different evidence from a 40% winner selected after thousands of attempts. Honest research preserves that context.
A disciplined TickRun workflow
- Begin with defaults. Run the selected strategy before changing parameters and compare it with Buy & Hold.
- Set realistic costs. Use an estimate appropriate to the type of trading being simulated.
- Inspect more than return. Review drawdown, Sharpe ratio, completed trades, individual trades, and the equity curve.
- Compare strategies once. Use the default comparison to understand the field without immediately optimizing everything.
- State the reason for optimization. Decide which parameter relationship you want to investigate and why.
- Limit the search. A larger maximum-combination setting increases exploration and the opportunity to fit noise.
- Stress the winner. Add costs, apply a signal delay, check nearby settings, and test other saved tickers.
- Treat it as a candidate. An optimized result is the beginning of validation, not the final conclusion.
Current TickRun limitation: the browser app tests the available saved ten-year range and does not yet provide train/test date splitting or walk-forward analysis. Comparing tickers and stressing assumptions is useful, but it is not a substitute for a genuinely unseen time period.
What optimization can—and cannot—tell you
Optimization is not inherently bad. It can reveal whether a strategy is insensitive to its settings, identify regions worth studying, expose interactions between parameters, and show that a plausible idea does not work under realistic assumptions.
What it cannot do is transform the best historical configuration into a forecast. The optimizer answers: “Which tested setup scored highest on this data under this model?” It does not answer: “Which setup will perform best next?”
The distinction is the foundation of responsible backtesting. Use optimization to generate questions, not certainty.
The practical takeaway
A good historical result is worth investigating, but its quality depends on the process that produced it. Ask how many alternatives were tried, whether the settings are stable, whether costs are realistic, whether enough trades occurred, and whether any truly unseen evidence exists.
The most trustworthy research is rarely the result with the highest number on the screen. It is the result that remains understandable and reasonably stable after deliberate attempts to break it.
Continue with the guide to all 40 TickRun strategies, read the user manual, or open the backtester.