azyware
Technology

Real-time ranking: personalising within a single session

EZ
Eazyware
· 7 min read
Quick answer

What should you know about real time personalisation?

Real-time ranking uses a feature store so a shopper's actions this session change what they see within seconds. Real time personalisation needs a streaming event path, an online feature store, a fast candidate generator and a lightweight ranker, with a nightly batch model kept underneath for the long-term profile.

Real time personalisation is the difference between a storefront that remembers what you bought last month and one that notices what you did thirty seconds ago. A visitor who has looked at three running shoes in a row should not see a homepage full of handbags on the next click. Delivering that needs more than a good model: it needs the click to reach a feature store within seconds, a candidate set retrieved in milliseconds, and a ranker light enough to run on every page view. This article explains the architecture, where the latency goes, and how to decide whether your business needs it or whether a nightly batch is enough.

Why in-session behaviour beats long-term history

Long-term profiles are valuable but slow to change, and they describe who a person usually is rather than what they want right now. Someone buying a gift for a friend looks, for that one session, nothing like their own history. Session based recommendations use the last few actions as the strongest signal, which is why they lift conversion even for visitors with rich profiles. For the many visitors with no history at all, the session is the only signal there is, which is the cold-start case in another guise.

Batch versus real time: what changes

ConcernNightly batchReal timeWhat it costs you
FreshnessRecommendations computed overnight from yesterday's dataFeatures updated within seconds of an eventStreaming infrastructure and an online store
Signals usedLong-term purchase and browse historyLong-term profile plus the current session's clicks, searches and cartTwo-tier feature design
ServingPrecomputed lists looked up by user IDCandidates retrieved and scored per requestA latency budget on the page
ModelLarge offline model, retrained dailySmall online ranker over batch-generated candidatesModel discipline: the ranker must be cheap
Failure modeStale but always availableFresh but must degrade gracefully when a component is slowFallback lists and timeouts
Right forEmail digests, weekly picks, slow-moving cataloguesStorefront, app home, search results, push at the right momentJudge by how fast intent changes in your business

The architecture of live ranking

Events in: the streaming path

Every view, search, add-to-cart and scroll event must leave the browser or app and arrive somewhere a model can read it within seconds. That is a stream, not a nightly export. The event schema, identity stitching and delivery guarantees are the topic of event pipelines; the short version is that real-time ranking is impossible on a pipeline that batches events hourly.

Feature store: the online and offline halves

A feature store keeps two copies of the same features. The offline store holds history for training and for batch precomputation; the online store, usually a low-latency key-value system, holds the current value per user and per item for serving. Session features such as "categories viewed in the last ten minutes" or "items in cart right now" are written to the online store by a stream processor as events arrive. The important discipline is that a feature is defined once and computed the same way in both halves, so the model sees at serving time exactly what it saw in training. Skew between the two is the most common cause of a real-time system that tests well and serves badly.

Candidate generation and ranking

You cannot score a catalogue of a hundred thousand items on every page view. Serving is two stages. Candidate generation pulls a few hundred plausible items quickly, from precomputed similar-item lists, from a vector search over item embeddings near the session's recent views, and from contextual popularity. The ranker then scores only those candidates with the fresh session features and the long-term profile, and returns the top slice. The ranker is deliberately small, often a gradient-boosted model or a compact neural network, because it has to run inside the page's latency budget.

The latency budget

A useful target is that the recommendation call adds nothing a visitor can perceive to page load. That means the whole path, feature lookup plus candidates plus ranking, has to fit inside a budget of a few tens of milliseconds at the high percentiles, not just on average. Set timeouts on every hop and fall back to a precomputed list when one is exceeded. A stale recommendation is far better than a slow page; conversion falls with every extra second of load, and no ranker is good enough to earn that back.

Feature store real time: the features that matter in a session

  • Categories, brands and price bands viewed in the last N minutes, with recency weighting
  • Search queries typed this session, including ones that returned nothing
  • Items added to or removed from the cart, which is a stronger signal than a view
  • Dwell time on product pages relative to the visitor's own norm
  • Entry point: campaign, landing page, referrer, device
  • Long-term profile features from the batch model, joined at serving time
  • Item-side freshness: stock level, price change, days since listing

Where real-time ranking applies beyond the storefront

The same machinery serves search result re-ranking, where the session context reorders results for a query; app home screens, where the tiles change as a user moves through the app; and message timing, where a push or WhatsApp nudge is sent only when session features suggest intent has stalled, as covered in personalised messaging. In B2B software the same design personalises a dashboard or a help panel based on what the user is doing right now, which is why we reuse it inside SaaS copilots.

Testing a real-time system honestly

Real-time ranking must be tested against the batch system it replaces, not against nothing, with users randomised once and revenue per visitor as the primary metric. Because the ranker changes quickly, an offline evaluation harness that replays logged sessions is essential: every candidate change or ranker retrain is scored on replayed sessions before it reaches traffic. This is the same evals-over-demos stance we take with every AI system, and it is described for personalisation in why you must run a controlled test.

A worked example

A retailer with a large catalogue and heavy paid traffic had a nightly batch recommender that worked for known customers and did nothing for the day's new arrivals, who made up most visits. We kept the batch model for the long-term profile and added a real-time tier on top: a streaming event path, an online feature store holding session features, vector-based candidate generation near the session's recent views, and a compact ranker with strict timeouts and a precomputed fallback. The homepage and category pages re-ranked on every view; search results were re-ordered using the same features. The system was launched behind a randomised test against the batch-only experience, and session-level lift was clear for both new and returning visitors. The retailer's merchandising team then asked for a rule layer so they could pin launches inside the ranked output, which we added with the ranker respecting the pins. The approach is what our personalisation engines service delivers as standard.

Team and timeline

A real-time tier needs a data engineer for the stream and feature store, an ML engineer for candidate generation and the ranker, a backend engineer for the serving API and fallbacks, and a product owner. On a clean event pipeline, a first placement goes live in about six weeks; where the pipeline needs work, add a phase for it. The build is scoped under personalisation engines, from $21,000 / ₹13.6L, and a serving API for your storefront or app is often paired with API development and integrations at $7,000 / ₹4.4L. Feature store and streaming infrastructure run on your cloud, owned by you, and monitored under the Care Plans listed on the pricing page.

Before you start: a checklist

  • Confirm events reach a stream within seconds rather than a nightly export
  • Agree the page latency budget and what happens when it is exceeded
  • Define each session feature once, with identical logic offline and online
  • Choose candidate sources: similar-item lists, vector search, contextual popularity
  • Keep the batch model for long-term profile features rather than replacing it
  • Build the session-replay evaluation harness before the first ranker change
  • Set up the randomised test against the current experience before launch
  • Decide how merchandising pins and business rules interact with the ranker

Glossary

  • Feature store: a system that stores model inputs, with an offline copy for training and an online copy for serving
  • Training-serving skew: a mismatch between features at training time and at serving time that degrades live results
  • Candidate generation: the fast first stage that narrows a catalogue to a few hundred plausible items
  • Ranker: the small model that scores candidates with fresh features and returns the top slice
  • Session feature: a value computed from the current visit, such as categories viewed in the last ten minutes
  • Fallback list: a precomputed recommendation served when a real-time component times out

Read recommendation engines explained for e-commerce leaders for the foundations and the D2C personalisation case study for an example across storefront and WhatsApp. Redis publishes documentation on using its data structures for real-time feature serving. Build and running costs are on the pricing page.

Real-time ranking is worth building when intent changes faster than your batch job runs; when it is, the feature store is the part to get right first.

Frequently asked questions

What is real time personalisation?

▾

It is ranking content or products using signals from the current session, updated within seconds of each action, rather than only a profile computed overnight. It relies on a streaming event path, an online feature store and a fast ranker.

Do we need a feature store for session based recommendations?

▾

For anything beyond a demo, yes. The online store holds fresh session features for serving while the offline store holds the same features for training, and defining them once prevents the skew that makes live results worse than tests.

How fast does live ranking need to be?

▾

Fast enough that a visitor cannot perceive it: a few tens of milliseconds at the high percentiles, with timeouts and precomputed fallbacks. A personalisation engine that slows the page loses more than it gains.