1. What you will learn
By the end of this lesson you will be able to define a supervised learning problem on financial data without using information that would not have been available, explain why random train-test splits fail on time series, compute exactly which training samples must be purged when labels look forward, and describe why fitting a data transformation on the whole dataset invalidates a result. This is education about research method and is not advice about trading or about any model.
2. The idea explained
A supervised model predicts a target from features. In finance the discipline lies entirely in the word available: every feature must be computable from information that existed at the moment the prediction is made, and the target must lie strictly in the future relative to that moment.
The first source is the label horizon. If the target is the return over the next five days, then a sample dated at some day carries an outcome that spans the following five days. A training sample close to the start of a test period therefore has an outcome that overlaps the test period, so the model is partly trained on the very interval it is being evaluated on. The remedy is purging: remove from the training set every sample whose label horizon reaches into the test period. Many practitioners also add an embargo, discarding a further short stretch after the boundary, because serial correlation extends the contamination a little beyond the arithmetic overlap.
The second source is the split itself. A random split assumes samples are independent, which financial data never are. Adjacent days are correlated, regimes persist for months, and a random split scatters samples from the same regime across both sets, so the model is effectively asked to interpolate within a period it has already seen. Evaluation must be chronological: train on earlier data, test on strictly later data, and report the dates.
The third source is preprocessing. Standardisation, imputation of missing values and feature selection all learn something from the data they are fitted to. If a scaler is fitted on the entire dataset, the mean and standard deviation it uses contain information from the test period. Every transformation must be fitted on training data alone and then applied unchanged to later data.
The fourth source is the universe. If today's list of instruments is used to study the past, every company that failed or was delisted has silently been removed, and the study measures the behaviour of survivors. The universe must be reconstructed as it stood on each historical date.
Finally, execution delay. A signal computed from a day's closing data cannot be acted on at that close. Document the delay assumed between signal and execution, and make it pessimistic.
3. The market, the regulator and the rulebook
Research of this kind is an offline exercise, and it should stay that way unless and until it is separately authorised. In India, securities markets are regulated by the Securities and Exchange Board of India, the statutory securities regulator constituted by an Act of Parliament, which sets the framework for algorithmic trading, for the approval and supervision of automated order flow through brokers, and for the registration of research analysts and investment advisers.
The exchanges publish the data definitions a study depends on: corporate action adjustments, index constituent changes with their effective dates, trading holidays and session timings.
Never state a brokerage rate, transaction charge, tax rate or margin figure as a fixed fact in a study. Assume each, label it, name the page publishing the current one with a date, and stress the assumption.
Worked example
4. Worked example
All figures are invented. A dataset has one thousand two hundred daily samples and the label is the return over the following five days, so a sample dated at day t carries an outcome spanning days t plus one through t plus five.
Split chronologically: training on days one to one thousand, testing on days one thousand and one to one thousand two hundred.
Now compute the purge exactly. A training sample at day t overlaps the test period when t plus five is at least one thousand and one, which means t is at least nine hundred and ninety-six. So the samples at days nine hundred and ninety-six through one thousand must be removed: five samples. Training therefore uses nine hundred and ninety-five samples, not one thousand.
Now the scaler. Suppose the correct training data, days one to nine hundred and ninety-five, has a mean of one hundred and a standard deviation of twelve. Suppose the full dataset, including the test period, has a mean of one hundred and eight and a standard deviation of eighteen, because the later period was more volatile.
Take a test observation of one hundred and forty. Standardised correctly, it is one hundred and forty less one hundred, divided by twelve, which is about three point three three: a genuinely extreme value that the model has never seen anything like. Standardised with the leaked scaler, it is one hundred and forty less one hundred and eight, divided by eighteen, which is about one point seven eight: an unremarkable value.
5. Common mistakes and how to fix them
The first mistake is a random train-test split. Split chronologically and report the exact date boundaries.
The second is forgetting to purge. Compute the label horizon, remove every training sample whose horizon reaches the test period, and state how many were removed.
The third is fitting transformations on all the data. Fit on training data alone and apply the fitted transformation unchanged to later data.
The fourth is using today's instrument list for a historical study. Reconstruct the universe as of each date, including delisted and failed names.
The fifth is assuming instant execution. Document a pessimistic delay between signal and execution, and test whether the result survives it. Most active traders lose money, past performance does not indicate future results, no model output is safe or guaranteed, and a SEBI-registered investment adviser is the person to consult about an individual's own money.
Key takeaways
6. Board summary
Every feature must be computable from information available at the prediction moment, and the label must lie strictly in the future. A forward-looking label overlaps the test period, so training samples within the horizon of the boundary are purged, often with an embargo. Financial samples are not independent, so evaluation is chronological with reported date boundaries. Scalers, imputers and feature selectors are fitted on training data only and applied unchanged afterwards. The historical universe must be reconstructed as it stood, and execution delay must be documented pessimistically.
Check your understanding
7. Practice and self-check
One. The label is a three-day forward return and the test set begins at day five hundred. From which day must training samples be purged? Day four hundred and ninety-seven onward.
Two. How many samples is that? Four, being days four hundred and ninety-seven to five hundred.
Three. Why not just the sample at day four hundred and ninety-nine? Because a sample at day four hundred and ninety-seven has a label spanning days four hundred and ninety-eight to five hundred, which already reaches the test period.
Four. What is an embargo, and why add one? A further stretch discarded after the boundary, because serial correlation extends contamination beyond the arithmetic overlap.
Five. A scaler fitted on all data gives a mean of fifty and a standard deviation of ten; fitted on training data alone it gives forty and five. A test value of sixty standardises to what under each? One point zero and four point zero.
Six. Which is correct and what does the other hide? The training-only figure; the other makes an extreme value look ordinary.
Seven. Why is a random split invalid here? Because adjacent samples and regimes are correlated, so both sets contain the same periods.
Eight. A study uses today's index constituents for a ten-year backtest. What has been removed? Every company that failed or was delisted, so it measures survivors.
Nine. What must be documented about execution? The assumed delay between signal and execution, chosen pessimistically.