Streaming, long jobs and background work in LLM apps
What should you know about LLM streaming UI and long-running AI jobs?
Stream tokens for interactive tasks, queue long jobs with progress, and never block a request on a slow model call. Those three rules decide whether an LLM feature feels responsive or broken. This article covers the streaming stack, the job queue, progress and cancellation, and what the build takes.
An LLM streaming UI shows the model's answer as it is generated instead of after it finishes; it is the difference between a feature that feels alive and one that appears to hang. But streaming only solves the first ten seconds. Model calls that take a minute, or a batch job that processes a thousand documents, need a different pattern: a queue, a job record, progress and a way to cancel. Every LLM application we ship uses both, and the choice between them is made per task, not per product.
This article sets out the three rules, the architecture behind each, the edge cases that catch teams out, and what it means for your build.
Why LLM streaming UI matters
Language models produce output token by token, and a full answer can take anywhere from two seconds to two minutes depending on length and model. A user staring at a spinner for eight seconds assumes something is wrong and clicks again, which doubles your cost and may create a duplicate action. Streaming shows the first words within a few hundred milliseconds, and people will read along with a stream far longer than they will wait for a blank screen.
Streaming also changes what the interface can do: show a plan before the answer, render citations as they arrive, and let the user stop a response that is heading the wrong way. None of that is possible if the response arrives as one block.
Three patterns, and which task gets which
| Pattern | Use when | How the user sees it | Typical duration |
|---|---|---|---|
| Streamed response | Interactive: chat, drafting, explanations, inline suggestions | Text appears as generated; stop button | Under 30 seconds |
| Queued job with progress | One long piece of work: a report, a large summary, a multi-step agent run | Job card with status, progress, cancel; notification when done | 30 seconds to minutes |
| Background batch | Many items: reprocessing documents, nightly enrichment, bulk classification | Not user-facing; admin view of throughput and failures | Minutes to hours |
The mistake is to pick one pattern for everything. Streaming a ninety-second agent run leaves a user watching intermediate noise; queuing a two-second reply makes a chat feel sluggish. Assign the pattern by the task's expected duration and whether the user needs to watch it happen.
The streaming stack
Transport
Server-sent events are the default. They run over ordinary HTTP, pass through most proxies and load balancers, reconnect automatically in the browser, and are enough for one-directional token delivery. WebSockets add bidirectional traffic, which matters for voice and for interfaces where the user interrupts constantly, at the cost of more infrastructure care. The Mozilla documentation on server-sent events is the reference.
What to stream besides tokens
A useful stream carries events, not only text: a status event when retrieval starts and ends, a tool-call event when the model decides to act, a citation event with the source, a final event with the complete answer and its metadata. The client renders each event type differently. This turns a wall of text into a readable account of what the system is doing.
The backend side
Provider SDKs from OpenAI and Anthropic stream tokens natively. Your service relays them to the client, but it also has to buffer the full response for logging and evals, apply any output filtering before tokens reach the user, and handle the client disconnecting mid-stream without leaking the model call. The last point matters for cost: cancel the upstream request when the user leaves.
Long-running LLM tasks: the job queue
Anything that takes longer than a request timeout, or that the user should not have to wait for, becomes a job. The user submits it, receives a job identifier immediately, and the work runs on a worker. This is ordinary background-job engineering, and the tooling is mature: a queue such as Redis-backed workers or a message broker, or a workflow engine such as Temporal when steps must survive restarts and retries.
A job record holds status, the current step, a progress estimate, partial results where they are meaningful, the error if any, and timestamps. The client polls or subscribes for updates. When the job finishes, the user gets a notification in the product and, if they have left, an email or push message. Never make the user keep a tab open to receive the result.
Progress that means something
Progress for LLM work is rarely a clean percentage, because you do not know how many tokens a step will produce. Report steps instead: "Reading 42 documents", "Summarising 12 of 42", "Drafting report". Users accept step-based progress readily and it is honest. Where a step has a countable inner loop, show the count.
Cancellation, retries and idempotency
Users will cancel, workers will crash and providers will return errors. Each job step should be idempotent, so a retry does not create a second draft or send a second message. Cancellation should stop the upstream model call and mark the job cancelled rather than leaving it running invisibly. Retries with backoff belong on provider errors, with a cap, and a job that exhausts retries should fail loudly with a readable reason.
Never block a request on a slow model call
This rule sounds obvious and is broken constantly. A request handler that calls a model synchronously and waits ties up a server thread or connection for the duration. Under load, a handful of slow model responses exhaust the pool and every other request, including ones that do not touch AI, starts failing. The symptom is an outage that appears to be caused by the database when it is actually caused by one prompt.
The fix is architectural: interactive calls stream from an async handler with a strict timeout and a fallback message; anything longer goes to a job. Set per-task latency budgets and enforce them. A model that has not produced its first token within a couple of seconds should trigger a fallback to a faster model or a graceful message, not an indefinite wait. This is one of the reasons our multi-model routing layer carries a fallback per task.
Observability across all three patterns
Every stream and every job should produce a trace: input, retrieved context, model, tokens, time to first token, total time, cost and outcome. For jobs, add per-step timing. Without this you cannot answer "why was it slow on Tuesday", and you cannot feed real cases back into the eval suite. LLM observability covers the tracing design in detail.
A worked example
A field-service SaaS company added a copilot with two very different jobs. The first was an inline assistant that answered questions about a ticket and drafted customer updates; the second was a weekly "account health" report that read every ticket, part order and site visit for a customer and produced a narrative with recommendations.
The inline assistant streams over server-sent events with status, citation and final events; users see the first words quickly and can stop a draft that is heading the wrong way. The report runs as a queued job on a workflow engine, with step-based progress ("Reading 118 tickets", "Summarising visits") and a notification when done. An early version had tried to run the report inside a request; under load it stalled unrelated pages. Separating the two patterns fixed both the experience and the reliability. The in-app copilot case study describes the product.
Team and timeline
Streaming and job infrastructure is built once and reused by every AI feature. A backend engineer and a front-end engineer typically need two to three weeks for the streaming transport, event types, job queue, progress UI, cancellation and tracing. It is part of the LLM application scope, starting at $21,000 / ₹13.6L, and of a SaaS copilot build from $19,500 / ₹12.8L. Where the work integrates with an existing backend and job system, the API development and integrations service covers the plumbing. Operating it after launch, including queue monitoring and provider incidents, fits a Care Plan. Prices are on the pricing page.
Before you start: a checklist
- Classify each AI task as interactive, long job or batch, with an expected duration
- Choose server-sent events for streaming unless you need bidirectional traffic
- Define stream event types: status, tool call, citation, final
- Cancel the upstream model call when the client disconnects
- Pick a queue or workflow engine and design idempotent steps
- Report progress as named steps, not fake percentages
- Set a per-task latency budget with a fallback model or message
- Trace every stream and job with tokens, timing and cost
Glossary
- Time to first token: delay between sending a request and the first streamed word appearing
- Server-sent events: a one-way HTTP stream from server to browser, well suited to token delivery
- Job record: the stored state of a background task: status, step, progress, result, error
- Idempotent step: a step that can be retried safely without repeating its side effects
- Workflow engine: infrastructure that runs multi-step jobs durably across retries and restarts
- Latency budget: the maximum time a task may take before a fallback is triggered
Related reading
What makes an LLM application production-ready, designing copilot UX and structured outputs and function calling cover neighbouring decisions. For durable job execution, the Temporal documentation is a good primary reference.
Stream what the user wants to watch, queue what they do not, and never let one slow model call take the rest of the product down with it.
Frequently asked questions
Should every LLM response be streamed?
▾
No. Stream interactive responses the user is watching, such as chat and drafting. Anything over roughly thirty seconds, or that the user need not watch, should run as a queued job with progress and a notification on completion.
Server-sent events or WebSockets for LLM streaming?
▾
Server-sent events for most text streaming: simpler, proxy-friendly and auto-reconnecting. WebSockets when traffic is bidirectional, such as voice or interfaces with constant interruption.
How do we show progress for a long AI job?
▾
Report named steps and counts, such as "Summarising 12 of 42 documents", rather than an invented percentage. Store progress in a job record the client can poll or subscribe to, and notify the user when it finishes.