azyware
Technology

Peak-load engineering for consumer apps in India

EZ
Eazyware
· 7 min read
Quick answer

What should you know about peak load architecture for a consumer app in India?

Design and load-test for month-end and festival peaks, not average days; queues, caching and autoscaling are non-negotiable. Indian consumer apps spike on salary days, sale events and match nights, so the architecture must absorb bursts, shed load gracefully and be rehearsed under test before the real day.

Peak load architecture for a consumer app in India starts from one fact: your average day does not matter. What matters is the evening of a festival sale, the first of the month when salaries land and bills are paid, the minute a match ends and everyone orders food, and the moment a marketing push notification goes out to your whole base at once. Those hours decide your reputation, your reviews and, for a marketplace or payments app, your revenue for the quarter. This article sets out how we design, test and operate consumer apps so that peak is a rehearsed event rather than an incident.

Why average-day sizing fails

Teams size infrastructure for the traffic they see in the dashboard, add a safety margin, and are surprised when a sale evening delivers many times that within a few minutes. The failure is rarely the web tier; it is a database connection pool that saturates, a third-party payment or SMS API that rate-limits you, a cache that expires all at once, or a single synchronous call in the checkout path that now takes seconds. Peak engineering is about finding those chokepoints before the day, and making the system degrade in a controlled way when one of them is hit anyway.

The peaks an Indian consumer app must plan for

PeakShapeWhat usually breaksDesign response
Festival and sale eventsSustained multiple of normal load for hours, sharp start at a published timeSearch, cart, inventory locks, payment gateway limitsPre-warm caches, queue orders, reserve stock asynchronously, agree gateway limits ahead
Month-start and salary daysPredictable daily spike in payments, bills and top-upsPayment webhooks, reconciliation jobs, bank-side latencyIdempotent webhooks, separate queues per provider, backlog-tolerant jobs
Live events and match endingsSudden burst measured in secondsAPI gateway, rate limits, hot keys in cacheEdge caching, request coalescing, per-user rate limits
Marketing pushesSelf-inflicted burst the moment a campaign sendsHome screen API, personalisation service, loginStagger sends, cache the landing content, load-test the campaign path
Regional outages and retriesRetry storms after a partial failureEverything, twiceExponential backoff with jitter in the client, circuit breakers on the server

Queues: turn bursts into backlogs

The single most important pattern is to accept work quickly and process it at a controlled rate. An order placement should write a small, durable record and return, with stock reservation, payment confirmation, notifications and downstream updates handled by workers reading from a queue. The user sees a confirmation with a pending state that resolves within seconds. When the burst exceeds worker capacity, the queue grows and latency rises gracefully instead of requests failing. Every consumer of the queue must be idempotent, because at-least-once delivery is what you get in practice, and the dead-letter queue must be watched by a person during the event.

Caching: protect the database from the crowd

Read traffic at peak is highly repetitive: the same home screen, the same category pages, the same few thousand products. Serve it from a cache in front of the application (a CDN for static and semi-static content) and a cache in front of the database (Redis or similar for hot objects). Two rules prevent the classic cache failure: expire keys with jitter so they do not all miss at once, and use request coalescing so that one miss triggers one database read rather than a thousand. Personalised content is cached per segment rather than per user where possible, and the personalisation service is designed to be skipped under load, returning a default ranking rather than an error.

Autoscaling: useful only if it is fast enough

Autoscaling on a cloud provider works for gradual growth over minutes and hours, not for a burst that arrives in thirty seconds. For a published sale start we scale ahead of time, based on the previous event and the marketing forecast, and let autoscaling handle the tail. Database capacity does not autoscale the way stateless services do, so read replicas, connection pooling and prepared capacity are decided in advance. Stateless services should start in seconds, which means small container images and health checks that pass quickly; a service that takes two minutes to become ready is not autoscalable in any practical sense.

Load-testing India apps: rehearse the real day

A load test that hits the home page with synthetic users tells you almost nothing. The test must replay realistic journeys, including search, add to cart, checkout with a sandbox payment, and the notification fan-out, at a multiple of the last peak, against an environment that matches production in database size and third-party integration. Run it at least twice: once weeks before the event to find chokepoints while there is time to fix them, and once a few days before to confirm the fixes and the scaling plan. Record the numbers; the next event's target is a multiple of this one's.

What to measure during the test

  • Latency percentiles (p95 and p99) per endpoint, not averages
  • Database connection usage and slowest queries under load
  • Queue depth and worker throughput; how long the backlog takes to drain
  • Third-party error and rate-limit responses from payment, SMS and maps providers
  • Cache hit ratio and the effect of a forced cache flush
  • Error budget: which requests failed, and whether users saw a useful message

Graceful degradation: decide what to switch off

Under sustained overload, something must give. Decide in advance what it is. Recommendations can fall back to a static list. Search can drop to a simpler ranking. Non-essential notifications can be delayed. A queue can display a waiting room for checkout rather than letting the database fall over. These are feature flags that an on-call engineer can flip in seconds, and they are tested during the load test, not invented during the incident. The point is to keep the payment path alive at the expense of everything else.

Operating the peak

The day itself needs a short runbook: who is on call, what dashboards are open, which flags exist, what the escalation path is to the payment gateway and cloud provider, and when scaling actions happen. A rehearsal call with the marketing team matters, because the single most common self-inflicted outage is a push notification sent to the entire base at one moment. Send in waves, and make sure the landing screen it points to is cached. Google's Site Reliability Engineering book is the best primary reference for the error budgets and incident practices behind this.

A worked example

A regional quick-commerce app came to us after a festival sale evening during which checkout timed out for most of the first hour. The store had been sized for the daily peak and the stock-reservation logic locked inventory rows synchronously during checkout. We rebuilt the order path around a queue with asynchronous reservation, moved catalogue and home-screen reads behind a cache with jittered expiry, added circuit breakers around the payment and SMS providers, and introduced flags to simplify search and delay marketing notifications under load. A replayed load test at a multiple of the failed evening's traffic was run twice before the next event. On the day, the queue backlog rose for a few minutes at the start and drained without user-visible failures, and the on-call engineer used one flag. The dispatch platform and driver app case study describes a different peak, the daily delivery wave, handled with the same patterns.

Team and timeline

Peak-readiness work on an existing app is typically a backend architect, a backend engineer for the queue and cache changes, a DevOps engineer for scaling and observability, and a QA engineer who owns the load tests, over four to eight weeks before the event. For a new build it is designed in from the start under our super app and marketplace development service from $63,000 / ₹41.6L, or SaaS development from $31,500 / ₹20.8L. Ongoing peak rehearsals and on-call cover sit under an Enterprise Care Plan at $5,250 / ₹3,40,000 a month, which includes 24×7 support and a named engineer. See the pricing page for all programmes.

Before you start: a checklist

  • List the peaks your app faces and the multiple of average each one represents
  • Map the synchronous calls in the checkout path and decide which move to a queue
  • Identify every third-party API in the critical path and its rate limits
  • Put a cache with jittered expiry and request coalescing in front of hot reads
  • Define the degradation flags and who may flip them
  • Build a realistic load-test journey and an environment that matches production data size
  • Agree the notification send strategy with marketing
  • Write the runbook and staff the on-call rota for the event

Glossary

  • p99 latency: the response time that 99% of requests beat; the number users on the slow tail experience
  • Request coalescing: merging simultaneous identical cache misses into one backend read
  • Circuit breaker: a component that stops calling a failing dependency for a period so it can recover
  • Backoff with jitter: client retries that wait increasingly longer plus a random delay, to avoid synchronised retry storms
  • Dead-letter queue: where messages go after repeated processing failures, for human review
  • Error budget: the amount of failure a service is allowed before feature work pauses for reliability

See super app architecture for the shared services that must scale together, wallets, UPI and payments in super apps for keeping the payment path idempotent under load, and peak-season readiness for online stores for the retail-specific version of this checklist.

Size for the worst hour, rehearse it twice, and decide in advance what you will switch off; that is the whole of peak-load engineering.

Frequently asked questions

How much headroom should we plan for a festival sale?

▾

Plan from your last measured peak and the marketing forecast, then load-test at a multiple of that. The exact multiple is specific to your app; the discipline of measuring and rehearsing is what matters.

Is autoscaling enough to handle peaks?

▾

No. Autoscaling handles gradual growth; bursts need pre-scaling, queues and caches. Databases and third-party APIs do not autoscale at all, so they need capacity agreed in advance.

How long does peak-readiness work take?

▾

Four to eight weeks on an existing app, including two load-test rounds. It sits under super app development or an Enterprise Care Plan; see the pricing page.