azyware
Technology

Event pipelines: the unglamorous foundation of personalisation

EZ
Eazyware
· 7 min read
Quick answer

What should you know about customer event pipeline?

A single event pipeline with a stable customer identity across storefront, app and messaging is what makes personalisation possible. A customer event pipeline defines the events once, stitches anonymous and known identities, delivers events within seconds, and keeps a replayable history for training and testing.

A customer event pipeline is the least exciting part of any personalisation project and the one that decides whether it works. Every recommender, next-best-action model and messaging decision is trained on and served from events: views, searches, carts, orders, opens, replies, support contacts. If those events are inconsistent across storefront, app and messaging, or if the same person appears as three different identities, the models learn noise and the measurements lie. This article covers what a good pipeline looks like, the identity problem that sits at its centre, and the order in which to build it so personalisation can start early rather than waiting for perfection.

Why the pipeline comes before the model

Teams often start with the recommender because it is the visible deliverable, then discover that the web tracking calls a product view one thing, the app calls it another, and WhatsApp interactions are not tracked at all. The model is trained on the web alone, serves everywhere, and the A/B test cannot assign users consistently because the app and web have different IDs. Fixing the pipeline afterwards means retraining and re-testing everything. Fixing it first costs a few weeks and makes every later model cheaper.

The components of a customer data pipeline

ComponentWhat it doesCommon failureWhat good looks like
Event schemaDefines each event name and its properties once, for every surfaceWeb, app and backend each invent their ownOne versioned schema, validated at ingestion
CollectionSDKs and server-side hooks that emit eventsClient-only tracking blocked by ad blockers and consentServer-side events for orders and messages; client for behaviour
Identity resolutionLinks anonymous sessions to a known customer across devices and channelsSeparate IDs per surface; no merge on loginA stable customer ID with a merge history
Streaming deliveryMoves events to consumers within secondsNightly exportsA stream with at-least-once delivery and idempotent consumers
StorageKeeps raw history for training and replayAggregates only; raw events discardedAppend-only raw log plus curated tables
Consent and retentionRecords what may be used for what, and for how longConsent captured on one surface, ignored on othersConsent as a property of the identity, enforced at read time

Designing the event schema

Start with the events the models will actually use and name them from the customer's point of view: product viewed, search performed, item added to cart, order placed, message sent, message opened, message replied, support contact opened. Each event carries a small set of required properties: timestamp, the customer or anonymous ID, the surface it came from, and the item or message it concerns. Keep the list short and version it. A schema with two hundred events, half of them unused, is harder to keep consistent than one with twenty that are validated on every write. Server-side events for anything that touches money or messaging are non-negotiable, because client-side tracking is lost to ad blockers and consent choices.

Identity resolution: the hard part

The problem

A visitor browses anonymously on a phone, later logs in on a laptop, receives an email, clicks through to the app, and replies on WhatsApp from a number that is not the one on their account. Personalisation needs all of that to be one person. Identity resolution is the process of stitching those threads: deterministic matching on login, email, phone and order details, and, where the business chooses, probabilistic matching on device and behaviour. We recommend deterministic matching only for most clients; probabilistic matching adds errors that are hard to explain to a customer who sees someone else's recommendations.

The mechanics

Every surface generates an anonymous ID for a new visitor. When a visitor identifies themselves, by logging in, entering an email at checkout or replying on a known number, the pipeline records a link between the anonymous ID and the customer ID, with the time and the reason. Events already recorded under the anonymous ID are then attributed to the customer for training, and future events flow directly. The link table is the most valuable asset in the pipeline, and it should be append-only so a wrong merge can be undone. Messaging channels need their own rule: a phone number that opts out on WhatsApp must propagate to the identity so email and push honour it too, as discussed in personalised messaging.

Streaming delivery for event tracking for personalisation

Batch pipelines that land events overnight are fine for reporting and useless for real-time ranking. Events should go onto a stream at collection time and be consumed by a stream processor that updates online features and by a sink that writes the raw log. At-least-once delivery is the sensible guarantee; consumers must therefore be idempotent, which means every event carries a unique ID and consumers ignore duplicates. Apache Kafka and its managed equivalents are the usual choice, and the Kafka documentation is the reference for delivery semantics. Ordering matters within a customer, so partition the stream by customer ID.

Storage: raw history you can replay

Keep the raw event log, append-only, for as long as your retention policy allows. It is the training set for every future model, the input for offline evaluation by replaying sessions, and the record that lets you rebuild a feature when its definition changes. Curated tables such as sessions, customer-day summaries and item statistics are derived from it and can be recomputed. The pattern is cheap in object storage and expensive to do without; teams that kept only aggregates find they cannot train the next model on the history they thought they had.

India's Digital Personal Data Protection Act and the GDPR both require that personal data is used for the purposes the customer was told about and retained no longer than necessary. In a pipeline that means consent is a property of the identity, captured on whichever surface the customer used, and checked when data is read for a purpose, not only when it is collected. Retention is a policy applied to the raw log per event type. Building this in from the start is easier than retrofitting it, and it is a topic in GDPR and AI systems. Our own handling of client data is described on the security page.

Build order: personalisation early, perfection later

  • Week one to two: agree the schema for the twenty core events and the identity rules
  • Week two to four: server-side events for orders and messages, client SDKs for browse behaviour, validation at ingestion
  • Week three to five: identity link table and login merge; anonymous history attribution
  • Week four to six: streaming delivery, raw log sink, first curated tables
  • From week six: first model trained on the raw log; online features fed from the stream
  • Ongoing: add events only when a model or report needs them; version the schema

A worked example

A D2C brand came to us for a recommender and a WhatsApp programme. The storefront had a marketing tag manager firing dozens of inconsistently named events, the app tracked a different set, and WhatsApp conversations lived only in the messaging provider. Orders were reliable because they came from the order system. We spent the first phase on the pipeline: a twenty-event schema, server-side order and message events, client tracking for browse behaviour, and an identity link table keyed on login, checkout email and verified phone number, with opt-outs propagated across channels. Events flowed onto a stream and into a raw log from the fourth week. The recommender and the messaging decision layer were then trained on a single, consistent history and served from the same identity, which is what made the later controlled test possible and the cross-channel results believable. The full engagement is described in the D2C personalisation case study.

Team and timeline

A pipeline build needs a data engineer for the schema, stream and storage, a backend engineer for server-side events and the identity service, a front-end or mobile engineer for client SDKs, and a client-side owner for consent and the definitions. Six weeks is typical for the build order above. The work is scoped as the first phase of personalisation engines, from $21,000 / ₹13.6L, or on its own under data and analytics applications, from $14,000 / ₹8.8L. The stream and storage run on your cloud and are owned by you. Care Plans on the pricing page cover schema changes, monitoring and the identity service.

Before you start: a checklist

  • List the twenty events the first models will use and name them from the customer's point of view
  • Inventory every surface (web, app, email, push, WhatsApp, support) and how each identifies a customer
  • Decide the deterministic identity rules and rule out probabilistic matching unless it is justified
  • Move order and message events server-side
  • Choose a streaming platform and confirm consumers will be idempotent
  • Define retention per event type and where consent is recorded
  • Plan validation at ingestion so a broken tag cannot poison the log
  • Agree who owns the schema and how changes are versioned

Glossary

  • Event: a timestamped record of something a customer did, with a small set of properties
  • Event schema: the agreed definition of each event and its properties, versioned
  • Identity resolution: linking anonymous and known identifiers into one stable customer ID
  • Link table: the append-only record of which identifiers were merged, when and why
  • At-least-once delivery: a guarantee that events arrive, possibly more than once, requiring idempotent consumers
  • Raw log: the append-only history of events kept for training and replay
  • Online feature: a value derived from the stream and served to models in real time

Continue with cold-start personalisation, which depends on tenure and item-age attributes flowing through the pipeline, and recommendation engines explained for e-commerce leaders. Services and prices are on the pricing page.

Build the pipeline first and every model after it gets cheaper, faster and more believable.

Frequently asked questions

What is a customer event pipeline?

▾

It is the system that collects customer behaviour as consistently defined events from every surface, stitches them under one identity, delivers them within seconds and stores a replayable history. It is the input to every personalisation model and measurement.

Why does identity resolution matter for personalisation?

▾

Without it the same person looks like several strangers across web, app and messaging, so models learn fragments and A/B tests cannot assign users consistently. Deterministic matching on login, email and phone fixes most of this.

How long does a customer data pipeline take to build?

▾

About six weeks for a twenty-event schema, server-side events, identity linking, streaming delivery and a raw log. It is delivered as the first phase of a personalisation engine so models can start in week six.