How long does LLM application development take? A realistic timeline
How long does LLM application development take?
LLM application development takes eight to twelve weeks for a single workflow and twelve to sixteen for two or three connected ones, including two to four weeks of shadow mode. Discovery adds ten days if the workflow is not yet chosen. Integrations and approvals, not model work, cause most slippage.
LLM application development takes eight to twelve weeks for a single workflow and twelve to sixteen weeks for two or three connected workflows, including two to four weeks of shadow mode before real users see output. Add ten days if the workflow has not been chosen yet. Integration access and approval cycles, not model work, cause most of the slippage.
Below is the week-by-week shape of a twelve-week build, the specific things that add weeks, what genuinely runs in parallel, and the shorter routes available when a date is fixed and the scope is not.
What actually consumes the calendar
People assume the long pole is getting the model to behave. It is not. On a twelve-week build, prompt work occupies perhaps ten days of engineering effort spread across the schedule. The calendar is consumed by four other things: agreeing what a correct output is, getting access to the systems the application must read and write, designing the interface humans use to accept or correct its work, and running it quietly in parallel with those humans long enough to trust the numbers.
That reordering has a practical consequence for planning. The dependencies that threaten your date sit inside your organisation rather than inside ours. A vendor can move fast on retrieval tuning and eval infrastructure; nobody can move fast on a security review that has not been booked or a set of credentials that needs a change advisory board.
It also means duration scales with breadth rather than with ambition. Asking for better answers on one workflow adds days. Asking for the same quality across three workflows, each with its own content, its own approver and its own integration surface, roughly triples the coordination and adds four to six weeks.
How long does it take, by scope?
Duration tracks the number of systems the application touches far more closely than it tracks the difficulty of the language task.
| Scope | Elapsed time | What ships | What you must supply |
|---|---|---|---|
| Discovery only | 10 days | Workflow choice, baseline, eval plan, cost model | Two workshops and access to reporting |
| Proof on the hard part | 3 weeks | Measured accuracy on the riskiest step | Sample data and a labelling reviewer |
| Single workflow, one system | 8 to 10 weeks | Production application, evals in CI, traces | API access, content owner, approver |
| Single workflow, multiple systems | 10 to 12 weeks | As above plus tool contracts and gates | Integration environments and credentials |
| Two or three connected workflows | 12 to 16 weeks | Shared retrieval layer, routing, approval flows | Cross-team product owner |
| Regulated or self-hosted deployment | 16 weeks and up | VPC or self-hosted inference, audit trail | Security review slot booked in advance |
Week by week: what a twelve-week build looks like
Weeks one and two: scope lock and the golden set
We freeze the workflow, record the current baseline from your own reporting, and start assembling one hundred to three hundred real inputs with human-labelled correct outputs. A reviewer on your side spends roughly a day a week on labelling in this period. This is also when refusal rules are written: what the application must never attempt, and what it must hand to a person.
Weeks three to six: retrieval, prompts and the first honest score
Content is ingested, chunked and indexed; hybrid search and reranking are tuned against the golden set; prompts are versioned in the repository and every change runs the suite. By the end of week six there is a number, and it is usually lower than the demo suggested. That is the point of running evaluation before polish.
Weeks seven to ten: tools, gates and the surrounding application
Each write the application performs becomes a narrow tool with a typed contract and a limit. Approval routing is built for anything above the thresholds your business owner signed. The user interface work happens here too: streaming, citations, confidence signalling and a clean hand-off path, which is more design work than teams expect. Teams that treat this as a thin wrapper over the model discover in shadow mode that reviewers cannot tell a good proposal from a bad one quickly enough to accept either, and the acceptance rate stalls for reasons that have nothing to do with answer quality.
Weeks eleven and twelve: shadow mode and instrumentation
The application runs alongside the people doing the work, producing output nobody downstream acts on, while the accepted-versus-corrected rate accumulates. Tracing and the cost dashboard go in during the same fortnight. OpenTelemetry publishes semantic conventions for generative AI spans, so this instrumentation follows an existing standard rather than being invented per project, which is one reason it fits in two weeks.
What adds weeks to an LLM application development timeline
Every delay we have seen in the last two years falls into one of these. None of them is about the model.
- Integration access. Waiting for a sandbox, credentials or a firewall rule is the single most common cause of a slipped date. Start the request in week one.
- Content that is not ready. Contradictory, undated or duplicated source documents force a cleanup that nobody scoped. Retrieval cannot resolve what people disagree about.
- No named approver. If the person who signs off thresholds and refusal rules is unavailable for a fortnight, the build waits, because the alternative is guessing at policy.
- Security review booked late. In banking, healthcare and education this reliably adds two to six weeks. Book the slot before the kick-off, not after the code is ready.
- Scope added mid-build. A second workflow is not a small change; it usually brings its own eval set, its own integrations and its own approver.
- Self-hosted inference. Procuring GPUs, sizing them and hardening the deployment adds three to six weeks against a hosted API.
- Labelling capacity. If the reviewer who labels the golden set is also running the operation, the set arrives late and everything downstream shifts.
What can genuinely run in parallel
Eval set creation runs alongside architecture work, because the two need different people. Interface design runs alongside retrieval tuning. Security review, once booked, runs alongside the build rather than after it, provided you can share the architecture document early. Content cleanup runs from week one to week eight in the background with your own team. Writing the runbook and the handover documentation belongs alongside the build rather than at the end, because the person who will own the system should be reading it while there is still time to argue with it.
What cannot be parallelised is the sequence of measure, change, measure again. Compressing that is how a project arrives on time with a system nobody trusts, and trust is the thing you cannot recover on a later sprint. If the date is immovable, cut scope rather than cutting the loop: one workflow delivered properly beats three delivered on faith, and the second and third are much faster once the retrieval layer, the eval harness and the approval pattern already exist.
What if you need something sooner?
Three shorter routes exist, and each answers a different question. A ten-day Sprint Zero through the AI Discovery Sprint at $3,250 or ₹2,00,000, credited to the build, answers what to build. A three-week ProofRun through the AI POC Sprint at $6,250 or ₹4,00,000 answers whether the hard step reaches usable accuracy on your data. A six-week Launch 6 MVP through AI-accelerated MVP from $26,500 or ₹17,60,000 puts a narrow version in front of real users.
The full LLM application development engagement runs from $21,000 or ₹13,60,000 to $84,000 or ₹56,00,000, fixed price against a locked scope, with the duration set by the tiers in the table above. All starting figures are on the pricing page.
When rushing the timeline is the wrong choice
If the date is driven by a conference rather than a customer, build the demo separately and keep it away from production data. Demos and applications have different acceptance criteria, and merging them produces a system with a demo's evaluation rigour and an application's blast radius.
If the compressed plan removes shadow mode, decline the compression. Shadow mode is the only period in which errors are free, and every week of it is worth two weeks of post-launch firefighting. If the compression removes the eval set, you no longer have a project, you have a prototype with a launch date attached.
What the schedule looked like in a real engagement
A B2B field-service SaaS company wanted an in-app copilot that could read customer data and draft actions against their existing API. Integration contracts and approval thresholds occupied more calendar than prompting did, and the rollout went cohort by cohort rather than all at once. The in-app copilot case study describes what shipped and in what order.
Checklist to protect the date
- Request integration credentials and sandbox access in week one
- Name the approver for thresholds and refusal rules before kick-off
- Book the security review slot at kick-off, not when code is ready
- Allocate a labelling reviewer for a day a week during weeks one to four
- Lock scope in writing and route additions to a second phase
- Keep two to four weeks of shadow mode in the plan whatever else moves
- Agree the go or no-go metric before shadow mode starts
Related reading
LLM application development: a practical implementation guide covers what happens inside each phase, what makes an LLM application production-ready sets the bar shadow mode is measured against, and questions to ask a vendor before you sign helps you test whether a quoted timeline is honest. To have a date sanity-checked against your scope, talk to us.
A timeline you can trust is one where the eval set and the shadow-mode window are the last things anyone is allowed to cut.
Frequently asked questions
How long does it take to build an LLM application?
▾
Eight to twelve weeks for a single workflow and twelve to sixteen weeks for two or three connected workflows, both including two to four weeks of shadow mode. Add ten days for discovery if the workflow is not chosen, and three to six weeks if inference must run on self-hosted infrastructure.
What is the fastest an LLM application can be delivered?
▾
A three-week ProofRun proves the hardest step on your own data, and a six-week Launch 6 MVP puts a narrow version in front of real users. Anything shorter is a demo, which is useful for a decision but should never be pointed at production data or live customers.
Why do LLM projects slip?
▾
Almost always for non-model reasons: waiting on integration credentials, source content that turns out to be contradictory, a late security review, or no named person to approve thresholds and refusal rules. Requesting access and booking the security slot in week one removes most of the risk.