Recommendation engines explained for e-commerce leaders
What should an e-commerce leader know about a recommendation engine?
A recommendation engine ranks products per shopper from behaviour and catalogue signals, with fallbacks for new users and products. It is layered: candidate generation, ranking, business rules and a controlled test that proves the lift, and it needs clean events and honest measurement more than a clever model.
A recommendation engine for ecommerce is the feature most often bought as a plugin and most often disappointing as one. The plugin shows "customers also bought" and nobody can say whether it earned anything. The version worth building ranks products per shopper from what they did, what the catalogue knows, and what the business needs to sell, with fallbacks for shoppers and products it has never seen, and with a controlled test that proves the lift. This article explains the parts, the decisions a leader has to make, the data it depends on, and what a build costs.
What a recommendation engine is and why it matters
At its simplest, the engine answers one question many times a second: for this shopper, on this page, right now, which products should appear in this slot? The inputs are the shopper's history (views, carts, purchases, searches), the session so far, the catalogue (categories, attributes, price, stock, margin), and what other shoppers did. The output is a ranked list, filtered by business rules, for a named slot: home page, product page, cart, search results, email or WhatsApp.
It matters because the slots already exist and are already showing something. The question is only whether that something is chosen well. Done properly, recommendations affect average order value, discovery of the long tail, and repeat purchase. Done badly, they show the shopper the thing they just bought. The personalisation engines practice at Eazyware builds the engine and, just as importantly, the measurement that says whether it worked.
The main approaches, compared
| Approach | How it works | Good for | Weakness |
|---|---|---|---|
| Popularity and rules | Best sellers, trending, merchandiser picks | Cold start, small catalogues, a baseline | Not personal; same for everyone |
| Item-to-item collaborative | Products bought or viewed together | Product page and cart slots | Needs traffic; ignores attributes |
| User-based collaborative | Shoppers similar to you liked these | Home page, email | Sparse for new or rare shoppers |
| Content-based | Similar attributes, text and images to what you engaged with | New products, niche catalogues | Over-narrow; needs good catalogue data |
| Learned ranking | A model scores candidates on many signals including session context | The final ordering across slots | Needs an event pipeline and a feedback loop |
| Session-based and real-time | Ranks within the current visit as it unfolds | Anonymous traffic, fast-moving intent | Engineering complexity, latency budget |
The architecture: candidates, ranking, rules
Production recommenders are layered. Candidate generation quickly produces a few hundred plausible products from several sources: items co-purchased with what is in the cart, items similar to recently viewed, items popular in the shopper's segment, items the merchandiser wants pushed. A ranking model then scores those candidates for this shopper and this slot using behaviour, catalogue and session features. Finally, business rules apply: in stock, not already purchased, margin floor, category diversity, no more than two from the same brand, exclusions for regulatory reasons. Each layer is testable on its own, and the rules layer is where the merchandising team keeps control.
Why the rules layer matters to leaders
The engine optimises what it is told to. Told to maximise clicks, it will show cheap impulse products. Told to maximise expected margin, it will hide the loss-leader that brings people back. The objective is a business decision, made explicitly and revisited, and the rules layer is where constraints the model cannot learn (stock, brand agreements, legal exclusions) are enforced.
Cold start: new shoppers and new products
Most traffic is anonymous or new, and the catalogue changes weekly. The engine needs fallbacks: for an unknown shopper, popularity within the landing category, then session-based ranking as soon as a few clicks arrive; for a new product, content-based similarity from its attributes, text and images, so it can be recommended before anyone has bought it. A recommender without cold-start handling performs well in the demo, on a shopper with a long history, and badly on the storefront, where most people have none. We cover the detail in cold-start personalisation.
The data it actually needs
Recommendations are only as good as the events behind them. The engine needs a reliable stream of product views, add-to-carts, purchases, searches and, ideally, impressions of the recommendations themselves so it can learn from what was shown but ignored. It needs a clean catalogue with consistent categories, attributes and stock. And it needs a stable identity for the shopper across sessions and devices where consent allows. In practice the first weeks of any build are spent on this pipeline, which is why event pipelines get their own article. If your storefront runs on Shopify, the Shopify developer documentation describes the storefront and webhook APIs the pipeline draws on.
Prove the lift or do not ship
The only honest measure of a recommender is a controlled test: a share of shoppers see the new engine, a share see the old slot, and the difference in revenue per visitor, conversion and average order value is measured over enough traffic and enough time to be meaningful. Click-through on the widget is not a business result. Neither is "revenue attributed to recommendations", which double counts what the shopper would have bought anyway. We build the test harness with the engine and run it before anything is declared a success; personalisation lift: why you must run a controlled test explains the design.
Beyond the storefront
The same ranked list feeds email, push and WhatsApp campaigns, where "products you might like" becomes a message rather than a slot. The constraints change: fewer items, a frequency cap, consent rules, and a stronger case for the merchandiser's rules. The personalisation and WhatsApp agent case study shows a D2C brand using one engine across the storefront and a conversational channel, and the retail industry page sets out where this fits among other retail AI work.
A worked example
A direct-to-consumer brand with a catalogue of a few thousand SKUs and frequent launches ran a plugin recommender that showed best sellers everywhere. Launch products never appeared because nobody had bought them yet, and repeat buyers saw items they already owned. The rebuild started with the event pipeline: view, cart, purchase and impression events from the storefront and the app, joined to a cleaned catalogue with attributes and imagery. Candidate generation combined co-purchase, attribute similarity (which solved the launch problem) and segment popularity; a ranking model scored candidates with session context; rules removed owned items and enforced category diversity. A controlled test ran on the product-page and cart slots first, with the old widget as control, and the merchandising team reviewed the top recommendations for a sample of shoppers each week during shadow mode. The engine then extended to the brand's WhatsApp channel, where the same candidates fed replenishment prompts. The measurable outcome was decided by the test, not by the widget's click rate, which is the point.
Team and timeline
A first recommendation engine is a data engineer for the pipeline, an ML engineer for candidates and ranking, and a product engineer for the slots and the test harness, over eight to twelve weeks. Pipeline and catalogue in the first three to four weeks; candidates, rules and a first slot in the next four; ranking model, controlled test and additional slots after. It is priced as a personalisation engine from $21,000 or ₹13.6L, and the AI/ML development service covers the ranking model where it is scoped separately. A ten-day Sprint Zero ($3,250, credited to the build) audits your events and catalogue first, which is usually where the risk is. Programs and Care Plans for ongoing retraining are on the pricing page.
Before you start: a checklist
- Audit the events you capture: views, carts, purchases, searches, impressions
- Check catalogue quality: categories, attributes, images, stock accuracy
- Decide the objective per slot: conversion, margin, discovery, repeat purchase
- List the business rules the engine must obey and who owns them
- Plan cold-start fallbacks for anonymous shoppers and new products
- Design the controlled test before the engine, with the old slot as control
- Confirm consent and identity rules for cross-session tracking
- Name the merchandiser who reviews recommendations weekly during shadow mode
Glossary
- Candidate generation: fast retrieval of a few hundred plausible products from several sources
- Ranking model: a model that scores candidates for a shopper and slot
- Collaborative filtering: recommending from what similar shoppers or co-purchased items suggest
- Content-based: recommending from product attributes, text and images
- Cold start: recommending for shoppers or products with no history
- Controlled test: an experiment comparing the new engine with the old slot on business outcomes
Questions clients ask
- Should we use a large language model for recommendations? Not for ranking. It is a machine-learning problem with a numeric feedback loop. Language models help with catalogue enrichment (attributes from descriptions and images) and with conversational discovery on top.
- Our catalogue is small. Is it worth it? With a few hundred SKUs, rules and popularity often suffice; personalisation earns its keep as catalogue and traffic grow. A Sprint Zero will tell you which side of that line you are on.
- Will it work on our app as well as the web? Yes, if events from both feed the same pipeline and the shopper identity is shared where consent allows.
- How often does the model retrain? Daily or weekly for the ranking model; real-time signals update within the session. Retraining runs under a Care Plan after launch.
- Do we own the model? Yes. Code, features, models, pipelines and documentation are yours at handover.
Related reading
Read real-time ranking: personalising within a single session, next-best-action personalisation and LLM or ML: choosing the right tool for why recommendation is still mostly a machine-learning problem rather than a language-model one.
Build the pipeline, layer candidates, ranking and rules, handle cold start, and let a controlled test, not the widget's click rate, decide whether the recommendation engine earned its place.
Frequently asked questions
How does an ecommerce recommendation engine work?
▾
It generates candidate products from co-purchase, similarity and popularity, ranks them for the shopper and slot with a model, and applies business rules such as stock, ownership and diversity before display.
Can a recommender handle new products and anonymous shoppers?
▾
Yes, with fallbacks: attribute and image similarity for new products, and category popularity followed by session-based ranking for shoppers with no history.
How much does a product recommendation system cost?
▾
A first engine with pipeline, ranking and a controlled test starts at $21,000 (₹13.6L) as a personalisation engine over eight to twelve weeks; see the pricing page.