A credible backtest must answer four questions: What was known? When was it known? At what price could a position change? What alternatives were tested before this result was selected?
1. Write the strategy before measuring it
Begin with a rule that another person could implement without asking what you meant. “Buy oversold stocks” is not a rule. A complete specification identifies the instrument, input fields, indicator formula, lookback, entry event, exit event, position size, execution delay, transaction costs, start and end conditions, and treatment of missing data.
Separate the economic hypothesis from the implementation. A trend hypothesis might claim that information is incorporated gradually. An implementation might buy when a 20-session SMA crosses above a 50-session SMA. The moving averages are not the reason the effect should exist; they are one measurement of the proposed effect. This distinction makes it easier to reject a weak implementation without endlessly tuning it.
Record the original specification before seeing results. Every result-driven change creates another tested alternative, even when the search is performed manually rather than by an optimizer.
2. Define the data contract
Backtesting requires more than a price column. You need to know whether prices are raw or adjusted, whether dividends and splits are represented consistently, which timezone defines a session, how duplicate dates are handled, whether volume is adjusted, and whether the history existed in the same form at the time.
TickRun reads saved daily OHLCV snapshots, validates their numeric fields, removes duplicate dates by keeping the last observation, sorts chronologically, and filters the requested interval. Its current files contain adjusted OHLCV history. That is appropriate for internally consistent indicator calculations, but it is not a complete point-in-time security-master database.
Testing AAPL today over its past ten years also contains a selection decision made today: you already know Apple survived and remained important. This does not invalidate a single-stock historical illustration, but it prevents the result from answering a universe-level question such as “Would this rule have selected successful stocks ten years ago?” That question requires point-in-time constituents including delisted securities.
3. Build an information timeline
The most important row in a backtest is often the boundary between signal and return. Suppose an indicator uses session t’s closing price. That closing value is not available before session t closes. Crediting the strategy with the return from t−1 to t would use information from the end of a return interval to earn that same interval.
TickRun prevents that particular look-ahead error with an explicit convention:
- Calculate the indicator and position state after close t.
- Record a buy or sell marker at close t.
- Apply the position held at t to the close-to-close return from t to t+1.
- Charge the configured position-change cost when the position changes.
This is an analytical convention, not a promise that an order can execute exactly at the marked close. A more detailed simulator would distinguish signal time, order-submission time, fill price, partial fills, and rejected orders.
4. Convert events into position state
Indicators normally produce candidate events, not a complete portfolio. A state machine must decide what happens when a buy arrives during an open position, when a sell arrives while flat, or when the history ends before an exit.
TickRun is long-only and all-in/all-out. A valid buy opens a position; subsequent buys are ignored until a valid sell closes it. Sells while flat are ignored. Signal delay shifts candidate events forward by a selected number of sessions. Minimum holding period suppresses an exit until enough sessions have elapsed.
An open position at the final date contributes to equity and drawdown but is absent from the completed-trade count and win rate. That difference is important: portfolio metrics use every daily return, while trade metrics use completed round trips only.
5. Model turnover and costs
Gross strategy return is not net investor return. A simulation should define commissions, bid–ask spread, slippage, fees, taxes where applicable, and market impact. These costs depend on instrument liquidity, order type, trade size, venue, and urgency; one universal percentage is not realistic.
TickRun uses a transparent first-order model. If the position changes from 0 to 1 or from 1 to 0, it subtracts the configured basis-point cost from that session’s strategy return. Ten basis points means 0.10% on entry and another 0.10% on exit. A completed round trip therefore incurs approximately 0.20% before considering compounding. Read the dedicated transaction-cost guide before treating that field as a complete execution model.
6. Select a benchmark that answers the question
A positive return is not enough. If the tested asset rose 150% while the strategy returned 40%, the strategy may have reduced risk, missed the market, or both. The benchmark supplies the counterfactual.
TickRun compounds Buy & Hold from the first to last available adjusted close using the same initial capital. It reports benchmark return and equity beside the strategy. The comparison is useful but incomplete: Buy & Hold stays fully invested, while a strategy may spend long periods in cash. TickRun currently assigns zero return to cash and does not credit a risk-free yield, so exposure differences affect the comparison.
7. Read a collection of metrics
Total return answers how wealth changed; it does not describe the path. Maximum drawdown identifies the worst observed peak-to-trough percentage decline. Sharpe ratio compares mean daily strategy return with its daily volatility and annualizes with √252; TickRun currently assumes a zero risk-free rate. Win rate counts positive completed trades, but says nothing about average win versus average loss.
Inspect the equity curve and trade list alongside summary numbers. A strong total return caused by one trade is different from a result distributed across regimes. A high win rate can coexist with poor performance when rare losses are much larger than frequent wins.
8. Separate development from evaluation
Parameters chosen on a history are fitted to that history. The highest-scoring configuration should be considered a candidate, not independent evidence. Reserve later data for a one-time test or use a chronological walk-forward process. Ordinary random train/test splitting is inappropriate when it allows future observations to inform a model evaluated on earlier observations.
TickRun currently evaluates and optimizes over the available saved ten-year range; it does not yet implement train/test partitions or walk-forward analysis. This means its optimized rankings are explicitly in-sample. The validation guide explains what a stronger workflow requires.
9. Preserve an audit trail
A reproducible research record should contain the data snapshot identifier, test dates, strategy version, parameters, cost assumptions, number of alternatives evaluated, code version, and complete metric output. Saving only the winning screenshot loses the search history needed to interpret it.
TickRun accounts save the ticker, tested dates, configuration, metrics, and completed trades. They do not save the full daily price/equity frame or the full optimizer search. For serious research, export or separately record those details until richer experiment tracking exists.
Backtest review checklist
- Was the rule specified before its result was viewed?
- Could every input have been known at the decision time?
- Is the fill convention later than the information used by the signal?
- Are adjusted prices, volume, splits, and dividends internally consistent?
- Does the universe include failures and delistings when the claim is cross-sectional?
- Are costs tied to turnover and plausible for the instrument?
- Is the benchmark comparable, and is idle cash handled explicitly?
- Do equity, drawdown, trade count, and individual trades agree?
- How many alternatives were tried?
- Was any period genuinely untouched until final evaluation?