azyware
Business

How much data do you need to train a model?

EZ
Eazyware
· 7 min read
Quick answer

How much data do you need to train a machine learning model?

It depends on the problem and the signal; assess in a discovery sprint and use pretrained or LLM approaches where data is thin. A strong signal with clean labels needs hundreds of examples, a weak one needs many thousands, and the honest way to find out is a learning curve on your own data rather than a rule of thumb.

How much data for machine learning is the first question every buyer asks and the one with the least satisfying general answer. It depends on how strongly the inputs predict the outcome, how clean the labels are, how many distinct situations the model must cover, and whether you are training from scratch or adapting something pretrained. What we can give you is a method: a way to find out, in a couple of weeks, whether the data you have is enough for the decision you want to automate, and what to do if it is not. That is what this article sets out.

How much data for machine learning, by problem type

The useful framing is signal-to-noise. If the outcome is nearly determined by a few inputs, as with a rule-like process that a person applies consistently, a model learns it from a few hundred examples. If the outcome depends on many weak factors and a lot of chance, as with churn, default or demand, the model needs enough examples for the weak signals to rise above the noise: thousands to tens of thousands of rows, and crucially, enough positive cases. A churn model with fifty churned customers in a hundred thousand does not have a hundred thousand examples; it has fifty.

ProblemRough training data requirementsWhat thins it outSmall-data alternative
Document field extractionA few hundred checked documents for evaluation; none for training with a vision-language modelLayout variety, handwriting, languagesPrompted general model with a review queue
Text classification (tickets, reviews, intents)A few hundred labelled examples per class to fine-tune a pretrained modelMany classes, ambiguous boundariesZero-shot LLM classification, then fine-tune on its corrected outputs
Tabular prediction (churn, default, fraud)Thousands of rows, with hundreds of positive cases at minimumRare outcomes, late labels, changing behaviourRules plus a simple model; collect labels deliberately
Demand forecastingTwo or more full seasonal cycles per seriesNew products, sparse SKUs, promotionsHierarchical models that borrow strength across similar series
Visual inspectionA few hundred labelled images per defect type to fine-tune a detectorRare defects, lighting and camera changesVision-language model for prototype and rare cases
RecommendationsMonths of interaction logs across a large user baseCold-start users and itemsContent-based and popularity baselines

Pretrained models changed the question

Most of the old training data requirements assumed you were teaching a model from nothing. That is rarely true now. For text, images and speech, a pretrained model already knows the language or the visual world, and you are teaching it your labels, which takes far less data. For many reading and classifying tasks, a large language model needs no training data at all: you describe the task, provide a handful of examples, and measure. The data you then need is for evaluation, not training, and a few hundred checked examples is enough to know whether the approach works. Small data ML today mostly means choosing the pretrained starting point well, as discussed in LLM or ML: choosing the right tool for the problem.

Where pretrained models do not help

Tabular business data is the exception. There is no pretrained model that knows your customers, your products or your transaction patterns. Churn, default, demand and fraud models are learned from your own history, and the data requirements above apply in full. This is where data readiness matters most and where a discovery sprint earns its fee.

Quality beats quantity, and labels are the constraint

A million rows with labels that were assigned inconsistently, or that leak the outcome, are worth less than ten thousand rows with labels that a domain expert would agree with. Before counting rows, we ask how the label was produced. Was it a system event (a chargeback, a cancellation) or a human judgement (a ticket category, a quality grade)? If human, did different people apply it the same way? If system, does the timestamp tell you when it became known, so that features can be computed as of the prediction date and leakage avoided? A short labelling exercise, where two experts label the same hundred examples independently, tells you how reliable your labels are. If they disagree often, no quantity of data will fix the model until the definition is fixed.

The learning curve: how to find out with your own data

The honest answer to "is it enough" comes from a learning curve. Train the model on a tenth of the data, then a fifth, then half, then all of it, evaluating each on the same held-out period. If the score is still climbing steeply at the full dataset, more data will help and it is worth waiting or collecting. If it has flattened, more of the same data will not help and the next gain comes from better features or a different approach. The scikit-learn documentation on learning curves describes the mechanics. We run this in the first two weeks of any tabular model project, and it converts the question from a debate into a chart.

ML data readiness: what we assess in a discovery sprint

  • Where the data lives, who owns it, and whether it can be extracted with history rather than as a current snapshot
  • How labels are produced, when they become known, and how consistent they are
  • How many positive cases exist, not just how many rows
  • Whether the features that matter are captured at the prediction moment or filled in afterwards
  • How much of the data is from the last year, since older patterns may no longer hold
  • Whether there are enough distinct entities (customers, stores, products) to generalise beyond the ones seen
  • What the model would replace, and what accuracy it needs to beat

What to do when data is thin

Thin data is a design constraint, not a stop sign. For text and image tasks, start with a general model and collect its corrected outputs as labels; within a few months you have the set to fine-tune something smaller and cheaper. For tabular tasks, ship rules that encode expert judgement first, log every decision and its outcome, and revisit the model when positive cases reach the hundreds. For forecasting, pool similar products or locations so the model learns shared patterns. In every case, the most valuable move is to start capturing the right data now: the decision, its timestamp, the inputs as they were, and the eventual outcome. Every month you wait to instrument is a month of training data you will never have. Our AI/ML development team designs that capture as part of the first build rather than after it.

A worked example

A D2C brand wanted a model to predict which first-time buyers would purchase again, so it could target retention offers. The order history was large, but the outcome was only reliable for customers whose first order was more than ninety days old, and the campaign data that would explain many repeat purchases was in a separate tool with no customer identifier. The discovery sprint produced a learning curve on the usable rows that flattened early: the available features did not carry much signal, however many rows were used. Rather than build a weak model, we recommended joining the campaign data, instrumenting the app to capture browsing signals, and running a simple segment-based rule for three months. The model built afterwards, with the new signals, was worth having. The personalisation work that followed is described in the WhatsApp personalisation case study.

Team and timeline

The data readiness assessment is the core of Sprint Zero: ten working days, $3,250 / ₹2,00,000, credited to the build, delivering a learning curve on your data, a label-quality check, a leakage review and a recommendation for the approach. It needs an ML engineer from our side and a data owner plus one domain expert from yours for a few hours a week. If the answer is "build", the model follows as an AI/ML development engagement from $17,500 / ₹11.2L. If the answer is "instrument first", the capture work often sits inside a data and analytics application build. All bands are on the pricing page.

Before you start: a checklist

  • Write the decision the model will make and when it will be made
  • Count positive cases, not rows
  • Find out how each label is produced and when it becomes known
  • Check that history can be extracted, not just a current snapshot
  • Have two experts label the same hundred examples and compare
  • List the features you believe matter and confirm they exist at the prediction moment
  • Decide what accuracy would make the model worth deploying, before seeing any results

Questions clients ask

  • Can we buy or synthesise data? External data can add features (weather, holidays, market prices). Synthetic labels for your own outcomes are rarely trustworthy; they encode whatever assumption generated them.
  • Is a year of history enough? For forecasting, two seasonal cycles are better. For churn or default, a year is often fine if positive cases number in the hundreds.
  • Should we wait until we have more data? Only if the learning curve is still climbing. Otherwise ship the rules-based version, instrument, and revisit.
  • Do LLMs need our data at all? For reading and classifying tasks, mainly for evaluation. For predicting your customers' behaviour, LLMs are not the tool.

Continue with MLOps for mid-size companies, demand forecasting with machine learning, and AI readiness assessment: the ten questions before you build.

Count positive cases, check the labels, draw the learning curve; then you will know, and nobody will have to guess.

Frequently asked questions

Is there a minimum number of rows for machine learning?

▾

No fixed number. What matters is positive cases, label quality and signal strength. A few hundred reliable positive examples can support a useful tabular model; a hundred thousand rows with fifty positives cannot.

Can we train a model with small data?

▾

For text and images, yes, by adapting a pretrained model or prompting a large one with a small evaluation set. For tabular predictions about your own customers, small data means shipping rules first and collecting labels deliberately.

How do we assess ML data readiness before committing budget?

▾

Run a discovery sprint: extract history, check how labels are produced, count positives, review for leakage and plot a learning curve. Ten working days gives a build-or-instrument answer with evidence.