azyware
Technology

Peak-season readiness: engineering your store for the sale

EZ
Eazyware
· 7 min read
Quick answer

What should you know about ecommerce peak season readiness?

Load-test to peak, cache aggressively, queue writes and rehearse the day; the sale is the design load. Peak-season readiness means treating the busiest hour of the year as the number your architecture is built for, proving it with a rehearsal weeks before, and having a runbook for the moment something still breaks.

Ecommerce peak season readiness is a simple idea that most stores get backwards: the sale is not an exception to normal traffic, it is the load the system exists to carry. Everything else is a quiet day. If your architecture was sized for the average Tuesday, the Big Billion Days, Diwali week or Black Friday will find the weakest component within minutes of the sale going live, and it will usually be the one nobody load-tested because it "never had a problem".

This article sets out how we prepare a store for a sale: what to test, what to cache, what to queue, how to rehearse the day, and who needs to be in the room. It applies whether you run a headless stack, a Shopify Plus store with custom apps, or a platform you built yourself.

Why the sale is the design load

Sale traffic is not just more of the same. It is spiky (a push notification at 12:00 lands tens of thousands of sessions in the same minute), it is write-heavy (carts, coupons, orders, payments, inventory decrements), and it is adversarial (bots scraping prices and hoarding stock). A system that serves catalogue pages comfortably at ten times normal traffic can still fall over at checkout at three times normal, because reads scale with caches and writes scale with locks.

So readiness cannot be a single number. You need a traffic model per path: landing pages, search, product pages, add-to-cart, coupons, checkout, payment callback, confirmation. Each has its own peak, bottleneck and fallback. The retail and e-commerce work we do starts by writing that model down before touching any code.

The readiness table: what to do for each path

PathTypical bottleneckWhat we do
Home, category, landing pagesOrigin servers rendering the same page millions of timesEdge cache with short TTLs, static generation for sale landing pages, cache keys that ignore tracking parameters
Search and filtersSearch cluster CPU, slow facet queriesPre-warm popular queries, cache facet counts, degrade to keyword-only search under load
Product pagesPrice and stock lookups on every renderCache the page, fetch price and stock as a small separate call, serve stale stock with a short refresh
Add to cart and couponsCoupon validation hitting the database per clickCoupon rules in memory, cart writes to a fast store, database write-behind
Checkout and inventoryRow locks on hot SKUs, oversellReserve stock via an atomic counter, queue order creation, confirm asynchronously
Payment callbacksGateway retries flooding the callback endpointIdempotent callback handler, queue and process, reconcile on schedule
NotificationsEmail and WhatsApp sends blocking the order pathEverything post-order goes through a queue with retries and rate limits

Load-test to peak, not to comfort

The most common load test we see is one that proves the system survives twice normal traffic. That is not a peak test. Take last year's busiest minute, multiply by your growth and by the marketing plan (a new celebrity campaign or a bigger push list changes the multiplier), and test to that number plus a safety margin. Test the full journey, not just page views: a script that browses, searches, adds to cart, applies a coupon and pays through the gateway's sandbox. Watch the database, the cache, the search cluster and the payment callback path, because the graph that goes vertical first tells you where the next week of engineering goes.

Run the test on production-shaped infrastructure; a staging environment with half the database and a smaller cache will pass and teach you nothing. Run it more than once, because the second run finds what was hiding behind the first bottleneck.

Cache aggressively and decide what may be stale

The single biggest lever for sale day scaling is deciding, explicitly, which data may be a few seconds old. Prices in a sale are fixed for hours; they can be cached at the edge. Stock is the hard case: shoppers hate "in stock" turning into "sold out" at checkout, but a stock call on every product view will melt the database. The usual answer is a cached stock indicator with a short TTL for display, and an authoritative atomic reservation at add-to-cart or checkout. The cache-control guidance on MDN is a good reference for getting edge and browser caching headers right; most stores we audit have at least one page that is uncacheable because of a stray cookie or a session-varying header.

Kill switches and degraded modes

Every non-essential feature needs a flag you can turn off in seconds: recommendations, reviews, live chat, personalised banners. Under load the store should degrade to a fast, plain version that still sells. Decide the switch-off order before the day and put the switches where the on-call engineer can reach them without a deploy.

Queue writes so the database never sees the spike

A high traffic store survives by putting a queue between the shopper and every slow or lockable write. Order creation, inventory decrement, loyalty points, invoice generation, notification sends, analytics events and CRM updates should all go through a queue with idempotent consumers, so a retry never creates a second order. The shopper sees an immediate "order placed" with a reference and gets confirmation a few seconds later. The gateway callback is the write that must be handled carefully: make it idempotent, accept it quickly, process it from the queue, and reconcile with the gateway's settlement report afterwards to catch anything dropped.

Hot SKUs deserve special treatment: a row lock per order on a flash-sale item will serialise checkout. An atomic counter in a fast store that reserves a unit before the order is created, with a timed release if payment fails, keeps checkout parallel and prevents oversell.

Rehearse the day

Black Friday readiness in India, where the season runs from Independence Day sales through Diwali and into the year-end, is a rehearsal, not a document. Two to three weeks before the sale, run a game day: replay the load test, deliberately fail a component (kill the search cluster, throttle the gateway sandbox, fill a queue) and watch the team respond. Time how long it takes to notice, to decide and to fix. Write down every gap in the runbook. Then freeze code changes for at least a week before the sale, so nothing new arrives untested. Bots and coupon-guessing scripts belong in the same rehearsal: rate-limit at the edge and keep coupon validation cheap; our note on fraud and anomaly detection covers card-testing.

The runbook

  • Who is on call, in which time zone, and how they are reached
  • Dashboards for each path with the thresholds that mean trouble
  • The order in which features are switched off under load
  • How to scale each tier by hand if autoscaling lags
  • How to pause marketing pushes if the store is struggling
  • How to reconcile orders and payments after an incident
  • Who talks to customers, and what the holding message says

A worked example

A D2C brand selling through its own store and WhatsApp had a sale go badly: the store stayed up, but coupon validation was slow, checkout timed out for some shoppers, and the team spent a week reconciling payments taken for orders never created. We rebuilt the coupon and checkout paths with in-memory rules, atomic stock reservation and queued order creation, made the payment callback idempotent, and moved every notification, including the WhatsApp agent messages, behind a queue. A load test to a multiple of the previous peak and a game day found the search cluster was the next weak point, so a keyword-only fallback was added. The following sale was, in the team's words, boring, which is the goal.

Team and timeline

A readiness engagement is typically an architect, a backend engineer and a performance engineer, with your platform owner involved throughout. Four to eight weeks is realistic: a week to model traffic and audit the paths, two to four weeks of fixes, a week of load testing and a game day, then a freeze. If checkout and order services need rebuilding for scale, it becomes a product engineering build; full-stack web work starts from $14,000 / ₹8.8L. For the sale itself, an Enterprise Care Plan gives 24×7 cover with a one-hour critical response and a named engineer who knows the system. Fixed prices are on the pricing page; a Sprint Zero produces the traffic model and prioritised fix list in ten working days.

Before you start: a checklist

  • Write the traffic model per path, from last year's peak, growth and the marketing plan
  • Audit which pages are actually cached at the edge, and why the rest are not
  • List every write on the order path and decide which go through a queue
  • Put atomic stock reservation on hot SKUs
  • Make the payment callback idempotent and add a reconciliation job
  • Build kill switches for every non-essential feature
  • Load-test the full journey on production-shaped infrastructure
  • Schedule a game day and a code freeze, and write the runbook

Glossary

  • Edge cache: a page or asset served from a CDN close to the shopper, so the origin never sees the request
  • Idempotent: an operation with the same effect whether run once or many times, essential for retries
  • Atomic reservation: decrementing stock in one uninterruptible step so two shoppers cannot both take the last unit
  • Game day: a rehearsal in which failures are injected on purpose to test the team and the runbook

Our guide to peak load engineering for consumer apps in India covers the app side of the same problem, demand forecasting for retail inventory helps you decide how much stock to reserve, and recommendation engines for e-commerce leaders explains what the personalisation layer should do when it is not switched off. The industries page for retail lists what else we build for stores.

Design for the busiest minute, prove it before the day, and the sale becomes a revenue event rather than an engineering one.

Frequently asked questions

How early should peak-season preparation start?

▾

Six to eight weeks before the sale if the store has not been load-tested before; three to four weeks for a store that has a runbook and only needs the numbers updated and a rehearsal.

Can a Shopify or headless store still fall over?

▾

Yes, through custom apps, coupon logic, third-party scripts and integrations that are not built for the spike. The platform scales; your customisations and back-office integrations are usually the weak point.

What is the single most valuable fix?

▾

Queueing writes on the order path with idempotent consumers. It protects the database, prevents duplicate orders on retry and makes recovery after an incident a reconciliation job rather than a manual clean-up.