azyware
Technology

Data leakage: the silent killer of ML projects

EZ
Eazyware
· 7 min read
Quick answer

What should you know about data leakage in machine learning and how do you prevent it?

Leakage lets future information into training so models look great and fail in production; validate on time, not random splits. It is the most common reason a model that scored well in a notebook disappoints in the first month live, and it is almost always a pipeline mistake rather than a modelling one.

Data leakage in machine learning is when the training data contains information that will not exist at the moment the model has to make a prediction. The model learns to use it, the validation score looks excellent because the validation set has the same information, and then the model goes live, the information is not there, and the score collapses. Nobody made a modelling error. Somebody made a pipeline error, and it was invisible because every metric said the model was fine. This article explains the forms leakage takes, why random train-test splits hide it, how time-based validation exposes it, and the review steps we run on every model before it ships.

What data leakage machine learning teams actually encounter

Textbook leakage is a column that literally encodes the label: a "refund issued" flag in a model meant to predict refunds. Real leakage is subtler and comes from how business systems store data. Records get updated in place, so the snapshot you train on reflects what was true after the outcome, not before. Aggregates are computed over the whole history, including the future relative to each row. Preprocessing is fitted on the full dataset before splitting. Each of these produces the same symptom: a validation score that is too good to be true and a production score that is merely true.

The main forms of leakage, and how each hides

FormExampleWhy random splits miss itFix
Target leakageA "churned" customer's status column is set before the prediction date in the training snapshotThe same column is in the validation setReconstruct features as of the prediction date
Temporal leakageRolling averages computed over all data including days after each rowRows from the future sit in both train and testPoint-in-time feature computation; time-based validation
Preprocessing leakageScaling, imputation or target encoding fitted on the whole datasetThe scaler has seen the test rowsFit transforms inside the training fold only
Duplicate leakageNear-identical records (same customer, same document) in both splitsModel memorises rather than generalisesSplit by entity, not by row
Label-proxy leakageNumber of collection calls made, used to predict defaultThe proxy is present at validation timeAsk when each feature is populated; drop those set after the decision
Selection leakageTraining only on applications that were approvedValidation set has the same selection biasModel the selection step or keep a bypass sample

Why random splits hide leakage

A random train-test split shuffles rows from every point in time into both sets. Any feature that carries information from the future is present on both sides, so the model gets to use it in validation exactly as it did in training. The metric measures how well the model exploits the leak, not how well it predicts. This is why so many ML validation mistakes look like success until launch. The scikit-learn documentation on common pitfalls covers the preprocessing case well; the temporal cases are the ones that bite business models hardest, because business data is almost always a time series pretending to be a table.

Time-based validation: the only honest test

The fix is to validate the way production will run. Pick a cutoff date. Train on everything before it, with every feature computed only from data that existed before it. Test on the period after it. Then move the cutoff forward and repeat, so you get a series of forward-window scores rather than one. The average across those windows is the number you report; the spread tells you how stable the model is as the world changes. If the forward-window score is far below the random-split score, you have found leakage, and now you can hunt for it. This is the validation we use in fraud detection, churn prediction and demand forecasting alike.

Point-in-time features

Time-based validation only works if the features are also computed as of the cutoff. A customer's "total orders" must be their total orders on the prediction date, not today. This usually means rebuilding features from event logs rather than reading them from a current-state table, which is more engineering than most teams expect and is the main reason a data engineer belongs on every AI/ML development project from day one.

Entity-level splits and duplicates

When the same customer, device or document appears many times, a row-level split puts some of their records in training and others in test. The model learns the entity rather than the pattern and scores well for the wrong reason. Split by entity so that everything about a customer lands on one side. For document models, near-duplicate scans of the same form are the classic case; deduplicate by content hash before splitting.

Overfitting in production: the symptom, not the cause

Teams often describe a model that fails after launch as overfitting production. Sometimes it is: too many features, too little data, a model that memorised noise. More often the model generalised perfectly well to the leaked signal and the signal disappeared. The distinction matters because the fixes are different. Overfitting is addressed with regularisation, simpler models and more data. Leakage is addressed by rebuilding the pipeline. Running the forward-window validation tells you which you have: an overfitted model scores poorly on both random and time splits; a leaky one scores well on random and poorly on time.

The leakage review we run before any model ships

  • For every feature, write down when it is populated in the source system and whether that is before the prediction moment
  • Recompute the top ten features by importance from raw events as of the cutoff and compare to the training values; any mismatch is a leak
  • Check that every transform (scaler, encoder, imputer) is fitted inside the training fold
  • Split by entity and by time, never by row alone
  • Compare random-split and forward-window scores; a large gap is a finding, not a curiosity
  • Look for features with suspiciously high importance; a single feature that dominates is usually the leak
  • Run the model in shadow mode against live traffic for at least one label cycle before trusting any number

A worked example

A subscription business asked us to review a churn model built in-house that scored very well in validation but had barely moved retention outcomes in three months of use. The pipeline read from the CRM's current-state customer table. One feature was "days since last login", computed on the day the training set was extracted rather than on the prediction date, so churned customers had large values and active ones small values, regardless of what was true when the retention team needed the prediction. Another was the count of support tickets, which included tickets raised during the cancellation process. Rebuilding both features from event logs as of each prediction date and validating on forward windows produced a much lower but honest score. The honest model was then improved with features that were genuinely available early. The retention team started acting on it because the list of at-risk customers finally matched their own experience. The same event-log discipline underpins our in-app copilot for a field-service SaaS, where usage signals feed both product analytics and models.

Team and timeline

A leakage review of an existing model takes an ML engineer and a data engineer about two weeks, and fits inside a Sprint Zero discovery sprint at $3,250 / ₹2,00,000, credited to any build that follows. Rebuilding features from event logs and re-validating is typically four to six weeks as part of an AI/ML development engagement, from $17,500 / ₹11.2L. Where the data foundation needs work first, our data and analytics applications team builds the event pipeline. Your side supplies someone who knows when each field in the source system gets written; that person is worth more than any algorithm. Price bands are on the pricing page.

Before you start: a checklist

  • Identify the prediction moment: exactly when will the model be called, and what data exists at that instant
  • List every source table and whether it is append-only or updated in place
  • Find the event log or audit trail from which point-in-time features can be rebuilt
  • Decide the entity to split on (customer, device, document, store)
  • Choose a cutoff date and confirm labels after it have had time to mature
  • Agree that forward-window score, not random-split score, is what gets reported
  • Plan a shadow period long enough for one full label cycle

Glossary

  • Leakage: information in training that will not be available at prediction time
  • Point-in-time feature: a feature computed using only data that existed at the prediction moment
  • Forward-window validation: training before a cutoff and testing after it, rolled forward repeatedly
  • Target encoding: replacing a category with a statistic of the label; a common leakage source when fitted outside the fold
  • Entity split: assigning all rows of one customer or document to a single side of the split
  • Shadow mode: running a model on live inputs without acting on its outputs, to compare against reality

Read MLOps for mid-size companies for what happens after validation, how much data do you need to train a model for the readiness side, and evals: the practice that separates AI demos from AI products for the LLM equivalent of this discipline.

If a model looks too good in the notebook, assume it is, and validate on time before anyone builds a plan around it.

Frequently asked questions

How do I know if my model has data leakage?

▾

Compare a random-split score with a forward-window, time-based score using point-in-time features. A large gap means leakage. A single feature with dominant importance, or a feature that is populated after the decision, is the usual culprit.

Is data leakage the same as overfitting?

▾

No. Overfitting is a model memorising noise and scores poorly on any honest test. Leakage is a pipeline fault that makes a model score well on flawed validation. The fixes differ: regularisation for one, pipeline rebuild for the other.

Does leakage apply to LLM applications too?

▾

Yes. Evaluation sets that overlap with few-shot examples or retrieved documents inflate scores in the same way. Keep eval questions out of prompts and indexes, and validate on questions written after the system was built.