Research principle: optimization proposes candidates. Stability around a candidate, economic logic, and performance on untouched observations provide the evidence.
A parameter grid is a response surface
Suppose an RSI strategy varies lookback, buy threshold, and sell threshold. Every combination maps to a backtest result. Together those results form a response surface. Reporting only the best cell discards the surface’s most informative feature: whether good performance occupies a broad region or one isolated spike.
A plateau—such as lookbacks 12–18 and buy thresholds 27–33 producing similar behavior—suggests the rule is not dependent on exact numerical luck. A lone winner at lookback 17 and threshold 29 surrounded by weak results demands skepticism. Markets do not usually contain a law that switches off when an arbitrary integer moves by one.
What TickRun’s optimizer searches
For the selected strategy, TickRun creates a deterministic grid over every relevant indicator parameter plus signal delays 0, 1, 2, and 3. Minimum holding period remains fixed at zero during optimization. Invalid window orderings—for example a fast period not shorter than a slow period—are skipped.
Each strategy exposes at least one million distinct raw configurations. The default maximum is 5,000 evaluations per strategy, so most runs sample rather than exhaust the space. Candidates are generated lazily. A deterministic coprime stride spreads early evaluations across dimensions, and the currently displayed configuration is evaluated first. The search stops on the combination limit, time limit, or complete valid grid.
The winner must satisfy the chosen minimum completed-trade count and is selected by compounded total return after configured transaction costs. It is not selected by Sharpe ratio or drawdown. “Best” therefore means highest return among eligible configurations actually evaluated under this dataset and budget—not globally optimal, robust, statistically significant, or likely to win later.
Search budget changes the experiment
A 5,000-combination run and a 100,000-combination run are not the same analysis. The larger search has more opportunities to find both genuinely useful regions and accidental winners. If two researchers use different time limits on different devices, they may evaluate different counts even with the same nominal grid.
Record evaluated count, stop reason, elapsed time, grid definition, and data snapshot. Do not compare winning scores from unequal search breadth as though selection pressure were equal. Increasing the budget after seeing an unsatisfactory result is an additional research decision and belongs in the trial count.
Test the winner’s neighborhood
After loading a winning configuration, vary one parameter at a time around it while keeping others fixed. Then inspect pairs of important parameters because interactions can be hidden in one-dimensional slices. For a fast/slow moving-average rule, useful coordinates include the fast window, slow window, their ratio, and absolute separation.
Compare more than final return. Record trade count, exposure, maximum drawdown, Sharpe, turnover, and the dates of major gains. Neighboring settings that produce similar total return from entirely different one-off trades are less reassuring than settings that preserve the same broad episodes.
Use relative perturbations where appropriate. Moving a 5-day window to 6 changes it by 20%; moving a 200-day window to 201 changes it by 0.5%. A fixed ±1 test therefore has different meaning across scales.
Stress non-indicator dimensions
Parameter stability includes assumptions that are not part of the indicator formula:
- Costs: increase basis-point friction across plausible scenarios.
- Timing: add one or more sessions of signal delay.
- Dates: shift start and end dates so one boundary event cannot dominate.
- Assets: apply the unchanged rule to instruments for which the hypothesis should plausibly hold.
- Regimes: separate trending, volatile, calm, and crisis intervals without selecting only favorable ones.
- Holding constraints: test whether rapid reversals create the apparent edge; TickRun requires this manual step because optimization fixes minimum hold at zero.
Minimum trade count is necessary, not sufficient
A configuration with two winning trades can report a strong return but provides little repetition. TickRun lets the researcher exclude candidates below a minimum completed-trade count. This prevents a low-trade setup from winning, but it does not establish a statistically adequate sample.
Trades overlap with market regimes, returns within a trade are serially structured, and repeated signals on one ticker are not independent coin flips. Fifty trades driven by one decade and one security do not automatically provide fifty independent observations. Inspect concentration: how much of final wealth came from the best trade, best year, or single crisis exit?
More trials raise the false-discovery problem
Testing parameters, strategies, tickers, cost assumptions, and date ranges creates a family of trials. Even if no repeatable edge exists, the maximum score tends to improve as that family grows. Research on multiple testing in finance shows why conventional significance thresholds cannot be interpreted independently after extensive data mining.
The visible optimizer count is only part of the multiplicity. Manual runs, discarded ideas, previous sessions, and alternative metrics also matter. A saved winner without its search history understates the selection process. Keep a research log that includes failures.
Choose from a plateau without inventing a new rule
If a broad stable region exists, prefer a simple, central configuration with economic meaning rather than the exact highest cell. The choice rule should be written before examining the final holdout. For example: select the median-performing configuration within the top stable cluster, subject to drawdown and trade-count constraints.
Do not average incompatible parameters blindly. A 10/50 moving-average pair and 50/10 pair are not symmetric because one violates the fast/slow relationship. Respect domain constraints and strategy semantics.
Freeze, then validate chronologically
Once the configuration and all decision rules are chosen, freeze them. Evaluate on later untouched data or inside a predeclared walk-forward process. If the result prompts another parameter change, that period becomes validation data, not a final test; obtain a new untouched period for a clean assessment.
TickRun currently optimizes the entire loaded interval and labels all-strategy optimized rankings “in sample.” It does not produce sensitivity heatmaps, cluster plateaus, adjust scores for multiple testing, or reserve a holdout automatically. Those steps remain the researcher’s responsibility.
Robustness checklist
- Was the hypothesis written before the search?
- Are grid boundaries and step sizes justified?
- How many total alternatives were considered?
- Did the run stop by limit, time, or exhaustion?
- Is the winner inside a broad stable region?
- Do neighboring settings preserve trades and risk, not just return?
- Does it survive costs, delays, date shifts, and relevant assets?
- Is performance concentrated in one observation or trade?
- Was the final rule frozen before untouched evaluation?