Start with a diagnosis
ZELVU Business logoZELVU BUSINESSBuild · Operate · Grow · Trade

Free sample lesson · no sign-up needed

Market Data Integrity: Timestamps, Corporate Actions and Survivorship Bias

From Algorithmic Trading, Python & Backtesting · Data integrity and causal signals · 10 min read

1. What you will learn

  • Why a backtest is only as trustworthy as the data it runs on, and what "data integrity" means for price and volume series.
  • How timestamps, time zones, trading sessions and bar conventions (open-high-low-close-volume) must be recorded and aligned.
  • How corporate actions such as splits, bonuses and dividends distort unadjusted prices, and how adjustment factors correct them.
  • What survivorship bias, look-ahead bias and data-snooping are, and how each one inflates backtested returns.
  • A repeatable checklist for auditing a dataset before any strategy code is written.

This lesson is exam and professional education using hypothetical examples. Nothing here is an investment recommendation or a promise of returns.

2. The idea explained

A backtest is an accounting simulation: it asks "if a fixed set of rules had been applied to past data, what trades, cash balances and positions would have resulted?" It is not proof that the rules will make money in future. Its value depends entirely on whether the past data is an honest picture of what a trader could actually have seen and traded at each moment. That honesty is called data integrity.

Point-in-time data. Point-in-time data is data recorded exactly as it was known on each historical date, without later revisions. Company financials are a classic trap: quarterly results are published weeks after the quarter ends and are sometimes restated. If a backtest uses the restated figure on the quarter-end date, it is using information nobody had then.

Timestamps and bar conventions. A daily bar labelled with a date usually summarises the whole session: the open is the first traded price, the close is the official closing price, and the high and low are extremes during the day. The key rule is that the close of a bar is only known after the bar ends. Intraday bars can be labelled by their start time or end time depending on the vendor, so a "09:15" five-minute bar may cover 09:15 to 09:20 or 09:10 to 09:15. Every dataset must document its convention, its time zone (Indian equity data is normally in IST, while many global feeds use UTC), and its session hours, including special sessions and holidays.

Missing values and bad ticks. Gaps occur because of holidays, trading halts, illiquid instruments or feed errors. Forward-filling a missing price is acceptable for valuing an open position, but it must never create a fake trade. A bad tick is an erroneous print (for example a price of 10 instead of 1,000); it should be flagged by a rule such as "a one-bar return beyond a set multiple of recent volatility with no matching volume", then checked against a second source.

Corporate actions. When a company splits its shares 1:5 (face value divided by five), the share price falls to roughly one-fifth overnight although no shareholder lost money. An unadjusted series shows a fake 80 per cent crash. A bonus issue of 1:1 (one new share for each share held) halves the price in the same way. Dividends cause a smaller drop on the ex-date. Data vendors produce an adjustment factor: historical prices before the event are multiplied by the factor so that the series is continuous. For a split of one share into n shares the price factor is 1/n; for a 1:1 bonus it is 1/2. Volumes are adjusted in the opposite direction (multiplied by n). Two kinds of series are common:

  • Price-adjusted (splits and bonuses only) — suitable for signals based on price levels.
  • Total-return adjusted (also dividends) — suitable for measuring what a holder actually earned.

A good practice is to store the raw series and the factor separately, so adjustments can be audited and re-applied.

Survivorship bias. Survivorship bias is the error of testing only on instruments that still exist today. Companies that were delisted, merged or went bankrupt disappear from many "current constituents" lists. A strategy tested on today's index members silently avoids every past failure, so its returns look better than they could have been. The cure is a point-in-time universe: for each historical date, use the list of stocks that were actually in the index, or actually listed and liquid, on that date, including those later delisted.

Look-ahead bias. Look-ahead bias is using any value at time t that became known only after t. Examples: trading at today's close on a signal computed from today's close; using an index membership list announced later; using a volatility estimate computed over the whole sample. The mechanical fix in code is to lag signals: the position for day t+1 is decided from information up to day t.

Data-snooping. Data-snooping (or selection bias) occurs when many rule variants are tried on the same history and the best one is reported. With enough attempts, something will look good by chance. We return to this in the walk-forward lesson; for now, note that a data log recording every experiment is part of integrity.

Licensing and provenance. Exchange data is licensed. A professional dataset records source, licence terms, download date, and any cleaning steps. This is not just legal hygiene: provenance lets another analyst reproduce the result.

3. Let us work through it

Step 1 — Document the source. Record the vendor, instrument identifiers (for Indian equities the ISIN is more stable than a ticker, because tickers can change after a name change), time zone, bar convention and licence.

Step 2 — Check the calendar. Compare the dates in the data with the exchange trading calendar. List missing sessions and unexpected extra rows (for example a weekend row caused by a time zone shift).

Step 3 — Validate each bar. For every row check low ≤ open ≤ high, low ≤ close ≤ high, volume ≥ 0, and no duplicate timestamps. Flag violations.

Step 4 — Detect jumps. Compute one-bar returns and flag any beyond a chosen threshold. For each flag, look for a corporate action on that date. If one exists, confirm the adjustment factor; if none exists, check a second source for a bad tick.

Step 5 — Build the point-in-time universe. For each date, list the instruments that were eligible then, including those later delisted. Store delisting dates and the final traded or settlement value.

Step 6 — Freeze and version. Save the cleaned dataset with a version number and a short change log. Every backtest result must cite the data version it used.

Worked example

4. Worked examples

Example 1 — Adjusting for a split. A hypothetical stock closes at ₹1,000 on the day before a 1:5 split and opens near ₹200 afterwards. Earlier closes were ₹950, ₹980 and ₹1,000.

  • Split factor for price = 1/5 = 0.2.
  • Adjusted earlier closes: 950 × 0.2 = ₹190; 980 × 0.2 = ₹196; 1,000 × 0.2 = ₹200.
  • The unadjusted one-day return would appear as (200 − 1,000) / 1,000 = −80 per cent. After adjustment it is (200 − 200) / 200 = 0 per cent, which reflects economic reality.
  • If a volume of 50,000 shares was recorded before the split, adjusted volume is 50,000 × 5 = 250,000.

Example 2 — A 1:1 bonus. A stock at ₹600 goes ex-bonus 1:1 and trades at about ₹300. Price factor = 1/2. A holder of 100 shares worth ₹60,000 now holds 200 shares worth about ₹60,000. A momentum rule on unadjusted data would wrongly register a 50 per cent fall and might trigger a sell.

Example 3 — Survivorship bias in a universe. A researcher tests "buy the ten cheapest stocks by price-to-book each year" using today's list of 50 index members. Suppose over the test period six companies that were in the index were later removed after severe losses. None of them appear in the test. Because cheap-looking stocks often include distressed firms, the missing names are exactly the ones this strategy would have bought. The backtested return is therefore biased upwards. The fix is to rebuild the index membership for each rebalancing date.

Example 4 — Detecting look-ahead in a column. A table holds close and signal, where signal on day t equals 1 if close on day t is above its 20-day average. The researcher computes strategy_return[t] = signal[t] × return[t], where return[t] is close-to-close from t−1 to t. Since signal[t] needs the close on day t, it was unknown during most of that return period. Correct version: strategy_return[t] = signal[t−1] × return[t] — in pandas terms, signal.shift(1) * ret.

5. Common mistakes and how to fix them

  • Using only today's index constituents for a ten-year test. Fix: build a point-in-time universe including delisted names and their final values.
  • Mixing adjusted and unadjusted prices in one calculation, such as adjusted closes with unadjusted stop-loss levels. Fix: choose one adjustment basis per calculation and store raw prices plus factors separately.
  • Assuming the vendor's time zone and bar label. Fix: read the documentation, then verify by comparing a known session open time with the first bar.
  • Forward-filling prices and then letting the strategy trade on the filled values. Fix: mark filled bars as non-tradable; use them only for valuation.
  • Deleting outliers automatically. Fix: flag, check against a second source or corporate-action list, and log the decision.
  • Not recording data versions. Fix: version every cleaned dataset and cite it in each result.

Key takeaways

6. Board summary

A backtest is an accounting simulation; its result is only as honest as its data. A bar's close is known only after the bar ends, so decisions must use lagged information. Splits and bonuses need price factors (1/n) and inverse volume factors (n). Survivorship bias comes from testing only survivors; use point-in-time universes. Look-ahead bias is using information not yet available; lag every signal. Store raw data, factors and versions so every result can be reproduced.

Check your understanding

7. Practice and self-check

  1. Define point-in-time data.

Answer: Data recorded as it was known on each historical date, without later revisions or restatements.

  1. A stock splits 1:2. What is the price adjustment factor for earlier prices?

Answer: 1/2, so each earlier price is halved.

  1. Before a 1:10 split the close was ₹2,500. What is the adjusted close?

Answer: 2,500 × 0.1 = ₹250.

  1. Why does survivorship bias usually inflate returns?

Answer: Failed or delisted companies are excluded, so the losses they would have caused never appear.

  1. Why is signal[t] × return[t] wrong when the signal uses the close of day t?

Answer: The signal was not known until the end of the period whose return it claims; it must be lagged one bar.

  1. What is the difference between price-adjusted and total-return adjusted series?

Answer: Price-adjusted corrects splits and bonuses; total-return also incorporates dividends.

  1. Name three bar-level validation checks.

Answer: Low ≤ open/close ≤ high, non-negative volume, and no duplicate timestamps.

  1. Why prefer ISIN to ticker for historical identification of Indian equities?

Answer: Tickers can change after renaming or restructuring, while the ISIN identifies the security more stably.

  1. What should happen to a forward-filled bar in a backtest?

Answer: It may be used to value an existing position but must not generate a trade.

  1. Why version datasets?

Answer: So every result can be traced to the exact data used and reproduced or audited later.

Found this useful?

7 more lessons like this are waiting.

Enrol to unlock the complete Algorithmic Trading, Python & Backtesting programme (4 modules · 8 lessons), the AI Business Tutor for questions about your own business, practice labs and 12 months of access.

One-time fee, paid upfront through ICICI Bank · bank financing assistance available · no subscription or auto-renewal.

Full programme

₹14,999

Enrol now