azyware
Business

How long does self-hosted AI agents take? A realistic timeline

EZ
Eazyware
· 7 min read
Quick answer

How long does self-hosted AI agents take?

Plan twelve to eighteen weeks from kick-off to an agent acting on its own: an eight to sixteen week build plus a shadow period. A single-task agent on cloud GPUs can be live in eight weeks. Owned hardware, several write integrations or a regulated review cycle push you to the top of the range.

Plan twelve to eighteen weeks from kick-off to an agent acting on its own, made up of an eight to sixteen week build plus a shadow period. A single-task agent served on cloud GPUs can be live in eight weeks. Owned hardware, three or more write integrations, or a regulated review cycle push you to the top of that range.

Below is the schedule as we actually run it on self-hosted agentic AI engagements: which workstreams overlap, which one is always on the critical path, and the five things that reliably add weeks.

What the clock is measuring

A self-hosted agent project has three finish lines and people quote whichever suits them. The first is a working agent in a development environment, which is fast and proves almost nothing. The second is the agent running in production alongside people, proposing actions that humans approve. The third is the agent completing a defined class of tasks without a human in the loop.

Only the third is a launch. When we say twelve to eighteen weeks, we mean to the third line, for one task family, with a widening plan for the next. Quotes that say six weeks are usually measuring to the first line, and the gap between the first and the third is where projects go quiet.

It helps to write the three lines into the plan explicitly, with a date and an owner against each. Steering committees respond badly to a project that appears to be ninety per cent finished for two months, and that is exactly what the stretch between a working demo and proven autonomy looks like from the outside. Naming the second line as a milestone in its own right, with its own success criteria, removes most of that friction.

The schedule, workstream by workstream

Elapsed time is shorter than the sum of these durations because most of them overlap. The column that matters is the last one.

WorkstreamTypical durationOverlaps withWhat blocks it
Discovery and scope lock10 working daysNothing, it gates everythingAccess to the people who do the task today
Infrastructure and model serving2 to 4 weeksTask design and eval authoringGPU procurement or a cloud quota increase
Evaluation set construction2 to 3 weeksInfrastructure buildAgreeing correct answers for real cases
Retrieval and indexing2 to 4 weeksTool contract workDocument access and permission metadata
Tool contracts and approval gates3 to 6 weeksRetrieval and indexingEach system owner signing off write scope
Agent build and hardening3 to 5 weeksLate retrieval tuningStable tools and a working evaluation loop
Shadow running3 to 6 weeksCare Plan onboardingReviewer availability, not engineering
Staged autonomy2 to 4 weeksNext task family designAcceptance rate holding at the agreed threshold

The critical path, and what is not on it

Infrastructure feels like the critical path and usually is not. Cloud GPU capacity can be running in days, subject to account quotas: accelerated instance types are capped per account by default and an increase has to be requested and approved, which is documented in the AWS Service Quotas guide and has equivalents at every major provider. Raise that request in week one, not week five.

The real critical path runs through permission to write. Every system the agent changes needs an owner who will agree what it may do, under what limits, with what approval. That conversation is organisational, not technical, and it moves at the speed of the people in it. On slow projects it is almost always the reason. Book those sign-off conversations in the first fortnight, with a draft tool contract in hand, because reviewing a concrete proposal takes an hour and defining one from scratch takes a month.

Evaluation set construction is the second constraint. Gathering one hundred to three hundred cases with agreed correct outcomes sounds administrative and turns out to be the hardest scheduling item, because it needs the time of the exact people who are busiest. Start it in week one and treat a missing eval set as a blocker, not a task. The useful framing for stakeholders is that the evaluation set is the specification: without it nobody can say whether the agent works, so there is nothing to sign off and no basis on which to widen autonomy.

What adds weeks

  • Owned hardware. Procurement, rack space and network change control add four to eight weeks that no amount of engineering effort compresses.
  • More than two write integrations. Each additional system adds a tool contract, a permission review and its own failure handling.
  • A regulated review cycle. Model risk review, security architecture sign-off and internal audit each run on their own calendar.
  • Undefined ground truth. If nobody can say what a correct outcome looks like, evaluation stalls and the build stalls behind it.
  • Legacy systems without APIs. Screen scraping or file drops are slower to build and far slower to make safe.
  • Changing the model family mid-build. Prompts, chunking and evaluation thresholds are all tuned to a model; swapping one costs a fortnight.

How long does the shadow period take?

Three to six weeks for a first task family, and it is the part to defend rather than compress. The agent proposes actions, reviewers accept or correct them, and every decision becomes evidence about where the system is reliable. You are not waiting for a date; you are waiting for an acceptance rate to hold steady at a threshold you set before launch.

Shadow length is driven by task volume, not by calendar. A process that runs two hundred times a day generates enough evidence in three weeks; one that runs twenty times a day needs longer. Shadow mode: the right way to launch AI agents in production describes how we move from proposing to acting, one action class at a time.

What the schedule costs

Eazyware delivers self-hosted agentic systems fixed price and fixed date, from $31,500 or ₹20,80,000 to $105,000 or ₹72,00,000 plus infrastructure. Before that, a ten-day Sprint Zero, sold as the AI discovery sprint at $3,250 or ₹2,00,000 and credited to the build, produces the task list, integration inventory and evaluation plan, which is what makes a fixed date possible. A three-week ProofRun at $6,250 or ₹4,00,000 tests the hardest task first. Every figure is on the pricing page.

After go-live, a Care Plan runs from $1,000 or ₹68,000 a month for Essential to $5,250 or ₹3,40,000 for Enterprise, with a $750 or ₹40,000 AI add-on covering evaluations and re-indexing. Book that onboarding during shadow running rather than after, so the first model upgrade is not also the first handover.

When speed is the wrong goal

If the deadline is external and immovable, self-hosting is usually the wrong shape of project. Capacity procurement, security review and shadow running are all poorly compressible, and the versions that go fast are the versions that skip the evidence. A hosted agent built on an API, moved in-house later, hits a hard date more reliably.

Compressing the shadow period is the most expensive saving available. Shipping an unproven agent into a process with write access does not save four weeks; it spends them on incident response, plus the credibility you needed for the second task family. If someone asks you to cut shadow running, the correct trade is narrowing scope to one action class instead. That genuinely saves weeks, because a narrower scope needs fewer tool contracts, fewer approval thresholds and a smaller evaluation set, and the remaining classes follow in a second phase at a fraction of the original cost.

What a real schedule looked like

A field-service SaaS company shipped an in-app copilot that could act inside its own product, described in the in-app copilot case study. The shape of that schedule is typical: tool contracts and approval thresholds took longer than the agent logic, and the system ran in shadow while dispatchers accepted or corrected its proposals before any action class went autonomous. The second and third capabilities landed far faster than the first, which is the pattern to expect: the first task family pays for the platform, and everything after it pays only for itself.

Checklist to protect the date

  • Raise cloud GPU quota or start hardware procurement in week one
  • Name one owner per system who can approve write scope, and book their time
  • Start evaluation set construction on day one, in parallel with everything
  • Lock scope after discovery and route changes through a written decision
  • Agree the acceptance threshold that ends shadow running before shadow running begins
  • Book security and model risk review slots early; they have their own queues
  • Keep staging on the same model version as production so upgrades are testable
  • Schedule the Care Plan handover inside the project, not after it

Self-hosted AI agents: a practical implementation guide covers the build stage by stage, evals over demos explains why the evaluation set gates the schedule, and five ways self-hosted AI agents projects fail covers the patterns that turn a fourteen-week plan into a thirty-week one.

The date you can hold is the one where discovery is finished, write permissions are agreed and the evaluation set exists; everything after that is engineering, and engineering is the predictable part.

Frequently asked questions

How long does it take to build a self-hosted AI agent?

▾

Eight to sixteen weeks to build, plus three to six weeks of shadow running before the agent acts alone, so twelve to eighteen weeks in total for a first task family. Owned hardware, several write integrations or a regulated review cycle push the schedule towards the upper end.

What usually delays a private AI agent project?

▾

Agreement on write permissions, not engineering. Every system the agent changes needs an owner to sign off scope, limits and approvals, and that runs at organisational speed. Building the evaluation set is the second constraint, because it needs time from the people who already do the work.

Can a self-hosted AI agent be delivered faster than twelve weeks?

▾

Yes, for one task family, one system, read access plus a single gated write, served on cloud GPUs with quota already in place. Eight weeks is achievable in that shape. Anything involving hardware procurement, multiple write integrations or formal model risk review will not compress that far.