Exception-handling agents for deliveries
What should you know about delivery exception management?
Exception agents read events, message the customer, reschedule within rules and escalate damage or refusal with context. Delivery exception management is the highest-return agent in last-mile logistics: each exception costs a rider's wait, a hub call and a customer's patience, and most follow a writable rule.
Delivery exception management is where a last-mile operation loses most of its time and most of its customers' goodwill. A rider arrives, the customer does not answer, the address does not resolve, the COD is not ready. The rider waits, calls the hub, the hub calls the customer, and ten minutes later somebody decides what to do. Multiply by every rider and every day. An exception agent removes the wait: it reads the event the moment the rider raises it, contacts the customer on WhatsApp, resolves the common cases within rules the operator has set, and hands the rest to a person with the whole story attached.
This article explains how a logistics agent for exceptions is designed: the event types and the policy for each, the conversation with the customer, the actions, the escalation path, the shadow-mode discipline, and what a build costs.
What a delivery exception agent is
It is an agent triggered by an event from the driver app or the tracking system, with read access to the order, the customer's contact preferences, the rider's location and the SLA, write access to a small set of actions gated by policy, and a channel to the customer, usually WhatsApp with SMS or voice fallback. It is not a chatbot the customer opens; it is a worker that opens the conversation when something goes wrong and closes it when the order has a new plan. See AI in logistics for the wider picture.
Exception types and the policy for each
| Exception | Agent may do alone | Agent proposes, human decides | Always human |
|---|---|---|---|
| Customer not available | Message, offer next slots within SLA, confirm, update route | Reschedule outside SLA | Third consecutive failure |
| Address not found | Send rider location, ask for landmark or pin, pass to rider, correct the order | Address change to a different pincode | Suspected fake address |
| Customer requests different slot | Reschedule within SLA and capacity | Same-day reschedule that breaks the route | None |
| COD not ready | Offer online payment link, offer next slot | Waive or reduce COD | Repeated non-payment pattern |
| Wrong item or quantity | Collect photo from rider, apologise, open return case | Partial delivery | Dispute on value |
| Damaged parcel | Collect photo, notify customer that a person will call, open case | Refund or replacement | Claim above threshold |
| Refused delivery | Collect reason, notify shipper, open return | Reattempt after shipper contact | Fraud indicators |
The table is the product. Every row is decided with operations and the customer-service lead before any prompt is written, and the agent's code enforces it. The policy-gated actions pattern is what makes the middle and right-hand columns safe: the agent can prepare a reschedule outside SLA with all the details, but a human clicks approve.
The conversation with the customer
The customer's side of a failed delivery automation should be short, honest and in their language. The first message names the order, the problem and the choices: "Your parcel from X could not be delivered because we could not reach you. Reply 1 for delivery tomorrow 10–1, 2 for tomorrow 2–6, or send your preferred time." Free-text replies are understood ("after 6 please", "leave with the security guard", a location pin), and anything the agent cannot map to an action goes to a person rather than a guess. Regional languages are handled by detection and reply, with slot names in the operator's standard form. WhatsApp template and session rules govern the first message and the reply window; the WhatsApp Business Platform documentation is the reference, and our post on WhatsApp AI chatbots for business covers the practicalities.
Address clarification
Address clarification AI is the single most valuable behaviour in Indian last-mile, because addresses are descriptive and pincodes are approximate. The agent sends the rider's position, asks for a landmark or a shared location, turns the reply into an instruction the rider can use, and writes the corrected address and geocode back to the order and the customer profile so the next delivery does not repeat the exercise. If the reply is ambiguous, it asks one clarifying question and then involves the rider or the hub.
Actions, written back as events
Every action the agent takes (reschedule, address update, payment link, case opened, shipper notified) is written to the order's timeline with the policy that permitted it and the customer message that triggered it. The driver app gets the update on its next sync; dispatch sees the changed plan; customer service sees the transcript if the customer calls. One timeline per order is what lets a human pick up an escalation without asking the customer to repeat anything, and it is the audit trail when a shipper asks why a parcel was rescheduled twice.
Escalation with context
Damage, refusal, suspected fraud, disputes and anything outside policy go to a person, and the hand-off is the difference between an agent that helps the hub and one that annoys it. The escalation carries the order, the exception, photos, transcript, the agent's proposed action and the reason it stopped, and routes to the right desk: hub, customer service or the shipper's contact. The weekly review of escalation reasons is where the policy table grows: a cluster of "leave with neighbour" requests becomes a new allowed action with its own rule. The same review discipline is described in how to build an AI agent that is safe to run unattended.
Evaluation and shadow mode
Before the agent messages a real customer, it runs in shadow mode: it reads real exception events and drafts the message and the action it would take, and the hub team compares the draft with what they actually did. Agreement per exception type is the acceptance criterion for switching that type to live; disagreement is analysed and either the policy or the prompt is changed. A golden set of a few hundred past exceptions with the correct outcome is built from this period and run on every change afterwards, so a prompt tweak that starts rescheduling outside SLA fails before release. The method is set out in shadow mode: the right way to launch AI agents.
Measuring the agent
- Exceptions resolved by the agent without human touch, per type
- Time from exception raised to new plan confirmed
- Successful delivery on the rescheduled attempt
- Rider wait time at the stop after raising an exception
- Actions outside policy (target zero) and escalations with complete context
- Customer satisfaction asked once after the exception is closed
- Address corrections that stuck, measured by the next delivery to the same customer
Report per exception type, never as an average; a high resolution rate on "customer not available" and a deliberate zero on "damaged parcel" is the correct shape. The argument for resolution over deflection is in AI ticket deflection is the wrong metric.
A worked example
A last-mile operator's hub teams spent most of each afternoon on the phone resolving failed attempts, and riders waited at stops for instructions. With the dispatch platform and offline-first driver app in place to supply events, we built an exception agent on WhatsApp with a policy table agreed in a two-day workshop with operations and customer service. It ran in shadow mode for three weeks while the hub team compared its drafts with their own decisions, and went live first for customer-not-available and address-not-found, then for slot changes and COD. Damage, refusal and anything outside SLA went to the hub with photos and transcript. Riders moved on as soon as they raised the exception, the hub's afternoon calls became a review of escalations, and corrected addresses accumulated against customer profiles.
Team and timeline
An exception agent is typically an AI engineer for the agent, policies and evaluation, an integration engineer for the driver app, dispatch and WhatsApp connections, and an operations owner from your side for the policy table and the shadow-mode review, over six to ten weeks: two weeks for the policy workshop, integrations and golden set, three to four weeks of build, and two to three weeks of shadow mode and staged go-live by exception type. It is scoped as a customer service agent build from $12,500 / ₹8L when the WhatsApp channel and a few actions are enough, or as part of a multi-agent systems build from $24,500 / ₹16L alongside dispatch. Running costs are model usage plus WhatsApp conversation fees, and a Standard Care Plan covers the weekly policy review. Current prices are on the pricing page, and the logistics sector page lists related work.
Before you start: a checklist
- Confirm exceptions arrive as events from the driver app or tracking system in near real time
- Rank exception types by volume from the last three months
- Fill in the policy table for each type: alone, propose, always human
- Define SLA windows, capacity limits and reschedule rules per customer segment
- Set up WhatsApp Business templates for the first message per exception type
- Decide the escalation desks and what context each needs
- Agree the shadow-mode period and the per-type acceptance criteria
- Name the operations owner for the weekly escalation review
Glossary
- Exception: any event that stops a delivery proceeding as planned
- Policy table: the agreed list of what the agent may do alone, propose or must escalate, per exception type
- Shadow mode: the agent drafting actions on real events while people still decide, for comparison
- Order timeline: the single log of events and actions on an order, readable by any team
- Session window: the period after a customer's message during which WhatsApp allows free-form replies
- Golden set: past exceptions with correct outcomes used to test every change
Related reading
See building an offline-first driver app for where the events come from, dispatch and routing engines for what happens after a reschedule, and returns and exchanges automation with policy-gated AI for the retail side of the same pattern.
Write the policy table with operations, run the agent in shadow, switch on one exception type at a time, and the rider stops waiting at the door.
Frequently asked questions
What delivery exceptions can an AI agent handle alone?
▾
Customer not available, address clarification, slot changes within SLA and COD-not-ready with a payment link, all within rules the operator sets. Damage, refusal, fraud indicators and anything outside policy go to a person with context.
How does the agent contact the customer?
▾
Usually WhatsApp, with SMS or a voice call as fallback, opened by the agent when an exception is raised. It understands free-text replies and location pins and asks one clarifying question before involving a person.
How long before the agent is live?
▾
Six to ten weeks including a shadow-mode period of two to three weeks, with go-live staged by exception type as each one meets its acceptance criterion.