azyware
Technology

How to build a support agent that knows the customer's order

EZ
Eazyware
· 6 min read
Quick answer

How do you build a support agent that knows the customer's order?

Account-aware support connects the agent to order, subscription and ticket data through permissioned APIs so every answer is specific: this order, this courier scan, this return window. Here is the architecture, the identity step, the integrations and the tests that keep it safe.

"Where is my order?" is the most common support question in commerce and the one generic bots answer worst, because they do not know which order. An account-aware agent does: it identifies the customer, reads the order and the courier's latest scan, and answers with the fact. This guide is the build recipe: identity, data access, tools, policy, testing and rollout, using a Shopify-style commerce stack as the running example, though the pattern applies to subscriptions, bookings and B2B accounts alike.

Architecture in one table

LayerWhat it doesExample
IdentityMatch the conversation to a customer, then verifyWhatsApp number → customer; confirm with order number or OTP
Read toolsFetch account data through permissioned APIsOrders, line items, fulfilment, courier tracking, subscription status, past tickets
KnowledgePolicies and help centre, indexed with citationsReturn window, exchange rules, delivery zones
Write tools (gated)Act within policyStart return, change address pre-dispatch, resend invoice, create ticket
EscalationHand off with contextHelpdesk ticket with order, transcript, summary
ObservabilityTrace every conversationIntent, tools called, outcome, cost

Step 1: identity before data

The agent must never read one customer's order to another. On WhatsApp the number is a strong first signal; on web chat a logged-in session is; on email the address is. A second factor, the order number, last four digits of the phone on the order, or an OTP, is required before any detail is shown. Identity is a tool with its own rules, enforced by the system rather than by the prompt, and it is the first thing tested with negative cases.

Step 2: read tools over your commerce stack

For Shopify, the Admin API exposes orders, fulfilments and customers; courier APIs give the latest scan; subscription apps give renewal status; the helpdesk gives past tickets. Each read is a small function with a stable output shape, cached briefly, and permissioned to the identified customer only. A custom OMS or ERP is wrapped the same way, sometimes as the first phase of a wider integration.

Step 3: knowledge with citations

Policies, help centre and resolved tickets are indexed with structural chunking, hybrid search and refresh on change, so the agent quotes the current return window rather than last season's. The retrieval discipline is in Why basic RAG fails in production; the point here is that knowledge and account data are combined in the answer: "Your order shipped Tuesday and is at the Bengaluru hub; our policy allows a return within 14 days of delivery, so you have until the 30th."

Step 4: write tools with policy gates

  • Start a return: only within the window, only for eligible items, only after confirming the reason
  • Change delivery address: only before dispatch, only within the same zone
  • Resend invoice, update contact details, apply an approved promo: low risk, allowed
  • Refund above a threshold, damage claims, disputes: create a ticket, never act
  • Every action previewed to the customer before execution and logged

Step 5: escalation that carries context

Escalations create a helpdesk ticket containing the customer, the order, the transcript, what the agent tried and a one-line summary of what is needed. The customer is told when to expect a reply. This is the difference between a hand-off and an abandonment, and it is how support teams come to trust the agent.

Step 6: tests that keep it safe

  • Negative identity tests: wrong number, wrong order, partial match must fail
  • Policy tests: every write tool tried outside its rules must refuse
  • Golden conversations: real chats with expected outcomes, run on every change
  • Prompt-injection tests: instructions in a customer message must not change behaviour
  • Load and failure tests: courier API down must produce an honest answer, not a guess

Step 7: rollout

Shadow mode for two weeks (agent drafts, team sends), assisted mode for two (agent acts with approval), then autonomous per intent as the record earns it. Weekly review of escalation reasons and unanswered questions. This sequence is described in AI customer service agents: resolve, don't deflect.

Beyond commerce

The same recipe serves subscriptions (plan, renewal, invoices), bookings (appointments, slots), banking (balances, cards, with stricter identity) and B2B accounts (contracts, tickets, usage). What changes is the list of read and write tools; the identity step, policy gates and escalation are constant. For B2B SaaS the agent usually lives in-app as a copilot.

A worked example

An apparel brand on Shopify connected the agent to orders, fulfilments, a courier aggregator and its helpdesk. Identity was WhatsApp number plus order number for any action. Read tools answered status and stock; write tools handled returns within the window and pre-dispatch address changes; damage and disputes created tickets. Golden conversations were built from two months of chats and rerun on every prompt change. The agent went autonomous for status in week two and returns in week four; the D2C case study has the outcomes.

What it costs and how long

An account-aware support agent on one channel runs from about $12,500 (₹8 lakh) and three to six weeks, with an AI engineer, an integration engineer and a delivery lead; multi-channel and multi-language deployments scale from there. See the customer service agent page and pricing.

Before you start: a checklist

  • API access to orders, fulfilments, courier tracking and the helpdesk
  • The identity rule for each channel
  • Written return, exchange and address-change policies with thresholds
  • Two months of chats for golden conversations
  • The escalation inbox and response expectation
  • Languages and sample chats in each

Handling the courier layer

Courier data is the messiest input: aggregators normalise scans differently, statuses lag, and "out for delivery" can mean anything. The read tool normalises statuses into a small set the agent can speak plainly, includes the timestamp of the last scan, and states uncertainty honestly ("the last update was at 2pm from the Bengaluru hub"). When the courier API fails, the agent says the tracking is unavailable and offers to message when it updates, which is a write tool with a scheduled retry.

Subscriptions and renewals

For subscription products the read tools cover plan, renewal date, invoices and usage; write tools handle plan changes within limits, invoice resend and payment-method links. Cancellations are the judgement intent: the agent gathers the reason and escalates to a person who can offer a save. Identity for billing actions uses a second factor even in-app, because the cost of a wrong action is money. The same golden-conversation discipline applies, with billing scenarios from real tickets.

Glossary

  • Read tool: a permissioned function that fetches account data
  • Write tool: a gated function that changes something, with preview and log
  • Second factor: an extra identity check beyond the channel identifier
  • Golden conversation: a real chat with an expected outcome, used in evaluation
  • Prompt injection: instructions hidden in customer text; tested against explicitly
  • Idempotency: safe retries for returns, refunds and address changes

Mistakes we see

The recurring mistakes: reading account data before verifying identity; giving the agent broad API scopes because it was quicker; putting refund limits in the prompt; skipping negative tests; and treating courier or OMS outages as impossible. Each is a design decision that costs a day to make and weeks to fix later.

Questions clients ask

  • Can it see the customer's whole history? Only what the task needs, fetched per conversation; nothing is copied into a separate store.
  • What if two orders match? It asks the customer to confirm which, never guesses.
  • Can it help pre-purchase? Stock and product questions, yes; personalised recommendations are a separate engine.
  • How do refunds get approved? Within thresholds automatically with a log; above them a ticket to a person.
  • Does it work on email? Yes; email is a channel adapter with the same identity rule using verification links.

What good looks like after 90 days

A ninety-day review checks identity-failure and policy-violation counts (target zero), read-tool error rates, resolution by intent, and the reopen rate on actions the agent took. Stable numbers are the signal to add the next channel or the next write tool.

WhatsApp AI chatbot for business, AI ticket deflection is the wrong metric and the retail industry page.

Identity, permissioned reads, gated writes, escalation with context, tests that prove each. Build in that order and the agent answers the question customers actually asked.

For decision-makers: the identity step and the policy gates are the two things to inspect before launch. If both are enforced by the system rather than the prompt, the rest of the recipe follows safely.

Frequently asked questions

Does the agent need direct database access?

▾

No. It reads and writes through permissioned API functions scoped to the identified customer; direct access is never granted.

What if our order system is custom?

▾

We wrap it with the same read and write functions; the agent does not care what is behind them.

Can it handle refunds?

▾

Within thresholds you set, with a preview and a log; above them it creates a ticket for a person.