azyware
Business

Recommendation engines explained for e-commerce leaders

EZ
Eazyware
· 7 min read
Quick answer

What should an e-commerce leader know about a recommendation engine?

A recommendation engine ranks products per shopper from behaviour and catalogue signals, with fallbacks for new users and products. It is layered: candidate generation, ranking, business rules and a controlled test that proves the lift, and it needs clean events and honest measurement more than a clever model.

A recommendation engine for ecommerce is the feature most often bought as a plugin and most often disappointing as one. The plugin shows "customers also bought" and nobody can say whether it earned anything. The version worth building ranks products per shopper from what they did, what the catalogue knows, and what the business needs to sell, with fallbacks for shoppers and products it has never seen, and with a controlled test that proves the lift. This article explains the parts, the decisions a leader has to make, the data it depends on, and what a build costs.

What a recommendation engine is and why it matters

At its simplest, the engine answers one question many times a second: for this shopper, on this page, right now, which products should appear in this slot? The inputs are the shopper's history (views, carts, purchases, searches), the session so far, the catalogue (categories, attributes, price, stock, margin), and what other shoppers did. The output is a ranked list, filtered by business rules, for a named slot: home page, product page, cart, search results, email or WhatsApp.

It matters because the slots already exist and are already showing something. The question is only whether that something is chosen well. Done properly, recommendations affect average order value, discovery of the long tail, and repeat purchase. Done badly, they show the shopper the thing they just bought. The personalisation engines practice at Eazyware builds the engine and, just as importantly, the measurement that says whether it worked.

The main approaches, compared

ApproachHow it worksGood forWeakness
Popularity and rulesBest sellers, trending, merchandiser picksCold start, small catalogues, a baselineNot personal; same for everyone
Item-to-item collaborativeProducts bought or viewed togetherProduct page and cart slotsNeeds traffic; ignores attributes
User-based collaborativeShoppers similar to you liked theseHome page, emailSparse for new or rare shoppers
Content-basedSimilar attributes, text and images to what you engaged withNew products, niche cataloguesOver-narrow; needs good catalogue data
Learned rankingA model scores candidates on many signals including session contextThe final ordering across slotsNeeds an event pipeline and a feedback loop
Session-based and real-timeRanks within the current visit as it unfoldsAnonymous traffic, fast-moving intentEngineering complexity, latency budget

The architecture: candidates, ranking, rules

Production recommenders are layered. Candidate generation quickly produces a few hundred plausible products from several sources: items co-purchased with what is in the cart, items similar to recently viewed, items popular in the shopper's segment, items the merchandiser wants pushed. A ranking model then scores those candidates for this shopper and this slot using behaviour, catalogue and session features. Finally, business rules apply: in stock, not already purchased, margin floor, category diversity, no more than two from the same brand, exclusions for regulatory reasons. Each layer is testable on its own, and the rules layer is where the merchandising team keeps control.

Why the rules layer matters to leaders

The engine optimises what it is told to. Told to maximise clicks, it will show cheap impulse products. Told to maximise expected margin, it will hide the loss-leader that brings people back. The objective is a business decision, made explicitly and revisited, and the rules layer is where constraints the model cannot learn (stock, brand agreements, legal exclusions) are enforced.

Cold start: new shoppers and new products

Most traffic is anonymous or new, and the catalogue changes weekly. The engine needs fallbacks: for an unknown shopper, popularity within the landing category, then session-based ranking as soon as a few clicks arrive; for a new product, content-based similarity from its attributes, text and images, so it can be recommended before anyone has bought it. A recommender without cold-start handling performs well in the demo, on a shopper with a long history, and badly on the storefront, where most people have none. We cover the detail in cold-start personalisation.

The data it actually needs

Recommendations are only as good as the events behind them. The engine needs a reliable stream of product views, add-to-carts, purchases, searches and, ideally, impressions of the recommendations themselves so it can learn from what was shown but ignored. It needs a clean catalogue with consistent categories, attributes and stock. And it needs a stable identity for the shopper across sessions and devices where consent allows. In practice the first weeks of any build are spent on this pipeline, which is why event pipelines get their own article. If your storefront runs on Shopify, the Shopify developer documentation describes the storefront and webhook APIs the pipeline draws on.

Prove the lift or do not ship

The only honest measure of a recommender is a controlled test: a share of shoppers see the new engine, a share see the old slot, and the difference in revenue per visitor, conversion and average order value is measured over enough traffic and enough time to be meaningful. Click-through on the widget is not a business result. Neither is "revenue attributed to recommendations", which double counts what the shopper would have bought anyway. We build the test harness with the engine and run it before anything is declared a success; personalisation lift: why you must run a controlled test explains the design.

Beyond the storefront

The same ranked list feeds email, push and WhatsApp campaigns, where "products you might like" becomes a message rather than a slot. The constraints change: fewer items, a frequency cap, consent rules, and a stronger case for the merchandiser's rules. The personalisation and WhatsApp agent case study shows a D2C brand using one engine across the storefront and a conversational channel, and the retail industry page sets out where this fits among other retail AI work.

A worked example

A direct-to-consumer brand with a catalogue of a few thousand SKUs and frequent launches ran a plugin recommender that showed best sellers everywhere. Launch products never appeared because nobody had bought them yet, and repeat buyers saw items they already owned. The rebuild started with the event pipeline: view, cart, purchase and impression events from the storefront and the app, joined to a cleaned catalogue with attributes and imagery. Candidate generation combined co-purchase, attribute similarity (which solved the launch problem) and segment popularity; a ranking model scored candidates with session context; rules removed owned items and enforced category diversity. A controlled test ran on the product-page and cart slots first, with the old widget as control, and the merchandising team reviewed the top recommendations for a sample of shoppers each week during shadow mode. The engine then extended to the brand's WhatsApp channel, where the same candidates fed replenishment prompts. The measurable outcome was decided by the test, not by the widget's click rate, which is the point.

Team and timeline

A first recommendation engine is a data engineer for the pipeline, an ML engineer for candidates and ranking, and a product engineer for the slots and the test harness, over eight to twelve weeks. Pipeline and catalogue in the first three to four weeks; candidates, rules and a first slot in the next four; ranking model, controlled test and additional slots after. It is priced as a personalisation engine from $21,000 or ₹13.6L, and the AI/ML development service covers the ranking model where it is scoped separately. A ten-day Sprint Zero ($3,250, credited to the build) audits your events and catalogue first, which is usually where the risk is. Programs and Care Plans for ongoing retraining are on the pricing page.

Before you start: a checklist

  • Audit the events you capture: views, carts, purchases, searches, impressions
  • Check catalogue quality: categories, attributes, images, stock accuracy
  • Decide the objective per slot: conversion, margin, discovery, repeat purchase
  • List the business rules the engine must obey and who owns them
  • Plan cold-start fallbacks for anonymous shoppers and new products
  • Design the controlled test before the engine, with the old slot as control
  • Confirm consent and identity rules for cross-session tracking
  • Name the merchandiser who reviews recommendations weekly during shadow mode

Glossary

  • Candidate generation: fast retrieval of a few hundred plausible products from several sources
  • Ranking model: a model that scores candidates for a shopper and slot
  • Collaborative filtering: recommending from what similar shoppers or co-purchased items suggest
  • Content-based: recommending from product attributes, text and images
  • Cold start: recommending for shoppers or products with no history
  • Controlled test: an experiment comparing the new engine with the old slot on business outcomes

Questions clients ask

  • Should we use a large language model for recommendations? Not for ranking. It is a machine-learning problem with a numeric feedback loop. Language models help with catalogue enrichment (attributes from descriptions and images) and with conversational discovery on top.
  • Our catalogue is small. Is it worth it? With a few hundred SKUs, rules and popularity often suffice; personalisation earns its keep as catalogue and traffic grow. A Sprint Zero will tell you which side of that line you are on.
  • Will it work on our app as well as the web? Yes, if events from both feed the same pipeline and the shopper identity is shared where consent allows.
  • How often does the model retrain? Daily or weekly for the ranking model; real-time signals update within the session. Retraining runs under a Care Plan after launch.
  • Do we own the model? Yes. Code, features, models, pipelines and documentation are yours at handover.

Read real-time ranking: personalising within a single session, next-best-action personalisation and LLM or ML: choosing the right tool for why recommendation is still mostly a machine-learning problem rather than a language-model one.

Build the pipeline, layer candidates, ranking and rules, handle cold start, and let a controlled test, not the widget's click rate, decide whether the recommendation engine earned its place.

Frequently asked questions

How does an ecommerce recommendation engine work?

▾

It generates candidate products from co-purchase, similarity and popularity, ranks them for the shopper and slot with a model, and applies business rules such as stock, ownership and diversity before display.

Can a recommender handle new products and anonymous shoppers?

▾

Yes, with fallbacks: attribute and image similarity for new products, and category popularity followed by session-based ranking for shoppers with no history.

How much does a product recommendation system cost?

▾

A first engine with pipeline, ranking and a controlled test starts at $21,000 (₹13.6L) as a personalisation engine over eight to twelve weeks; see the pricing page.