Questions to ask a recommendation engine development vendor before you sign
What should you ask a recommendation engine development vendor?
Ask a recommendation engine development vendor for a shipped system and the holdout result it produced, how they will handle your event data, who owns the model artefacts, what it costs to run for two years, and what happens on the day quality drops. The answers sort the field quickly.
Start with one question: describe a recommendation engine you shipped, the metric it moved, and how you measured it against a holdout. A vendor who has done this answers in specifics within two minutes. From there, ask about your event data, model and artefact ownership, the two-year running cost, the retraining cadence, and what happens the week quality drops.
This article groups the questions into five themes, gives the answer you want to hear and the warning sign beside it, and names the questions that sound rigorous but tell you nothing. It also covers when running a vendor process is premature.
The question that sorts the field fastest
Most personalisation pitches are demonstrations. A demo built on a sanitised slice of a catalogue proves that the vendor can rank items, which was never in doubt. What is in doubt is whether they can make a ranking survive real event data, real latency budgets and a real control group.
So ask for the evidence rather than the demo: a system in production, the surface it ran on, the primary metric, the holdout percentage, the test duration, and what the result was including the times it was flat. A vendor who has only ever reported offline accuracy has not been through the part of the project where the money is decided. Our own position on this is set out in evals over demos.
The structure of the rest of the conversation matters less than its coverage. NIST's AI Risk Management Framework organises practice around governing, mapping, measuring and managing, which is a serviceable spine for due diligence: who is accountable, what the system is for, how quality is measured, and what happens when it drifts.
Five themes, good answers and warning signs
| Theme | The question | Good answer | Warning sign |
|---|---|---|---|
| Evidence | Show a shipped engine and its holdout result | Named surface, metric, holdout size, duration, honest outcome | Offline accuracy only, or a demo dataset |
| Your data | What will you do about our event layer? | Audit first, server-side contract, identity stitching, priced separately | We can work with whatever you have |
| Approach | What will you build for our catalogue size? | Starts simple, adds complexity only when data supports it | Same architecture proposed to every client |
| Ownership | Who owns the code, pipelines and model artefacts? | You do, in your repositories and cloud accounts | Hosted black box with an export on request |
| Operating | Who retrains it and how often? | Named cadence, pipeline, rollback, in a support agreement | It learns automatically |
Questions about your data, not their model
Ask what they will do in the first fortnight. The answer should involve your event stream before it involves any modelling: which events exist, how they are collected, whether anonymous sessions are stitched to accounts, how much history is usable, and what has to be rebuilt. Any vendor who proposes to start modelling in week one has either audited your data already or is about to discover the problem at your expense.
Then ask a harder one: what would make you tell us not to build this? The useful answers are concrete. Too few monthly active users, too small or too static a catalogue, no ability to hold out traffic, no owner for the primary metric. A vendor with no disqualifying conditions has one product and is selling it regardless of fit.
Finally, ask how they handle the cold start for your catalogue specifically. If a meaningful share of your inventory is new each month, the fallback design is not a detail, it is most of the perceived quality in the first quarter.
Questions about ownership and exit
- Do we own the trained model artefacts, feature definitions and evaluation sets? The answer should be yes without qualification.
- Does everything run in our cloud accounts and our repositories? Hosted-only arrangements make the exit conversation expensive.
- Who holds the provider API contracts? Usage should be billed to you, with budgets and dashboards set during the build.
- What does handover include? A runbook, retraining instructions, dashboards and working sessions with your engineers, not a slide deck.
- How long does it take us to run this without you? Ask for weeks, and ask what the first thing to break would be.
- What are the notice terms and the exit deliverables? Agree them before signature, not during a dispute.
- Can we run our own evaluation before acceptance? A vendor confident in the number will say yes immediately.
Our position is that you own the code, the infrastructure, the prompts, the model choices and the documentation, and the reasoning behind it is in the glossary entry on code and IP ownership. Ask every bidder to match it in writing rather than in a meeting.
What should you ask about cost?
Ask for a twenty-four month total, not a build price. Split it into build, event remediation, embedding and inference usage, storage, experimentation, retraining and support. A vendor who will not produce that split either has not run a system for two years or is hoping you will not notice where the money goes after launch.
For reference, Eazyware builds recommendation and personalisation engines from $21,000 or ₹13,60,000 to $70,000 or ₹46,40,000, and post-launch cover runs from $1,000 or ₹68,000 per month on an Essential Care Plan to $5,250 or ₹3,40,000 on Enterprise with a named engineer, plus $750 or ₹40,000 for the AI system add-on covering evals, re-indexing and cost monitoring. Those figures sit on our pricing page and on the maintenance and support page so a buyer can compare without a call.
Ask also what the vendor does when the number comes back flat. The honest answer involves diagnosis and a decision point, not an automatic second phase. If every outcome leads to more work at your expense, the incentives are wrong.
Questions about the week after launch
Recommendation quality decays. Catalogues turn over, seasons change, and a promotion reshapes what people click. Ask who notices, how quickly, and through what alert. Ask what the rollback is and whether it has been tested. Ask how long a retrain takes end to end and who is allowed to approve one.
Then ask about the control group after go-live. Many teams dissolve the holdout once the engine ships, which means the next twelve months of changes are unmeasured. A vendor who insists on keeping a small permanent holdout is thinking about your second year, and personalisation lift, why you must run a controlled test explains why that discipline pays for itself.
Questions about how they will work with your team
Ask who is actually on the engagement and how much of their week you get. Named people with named hours is a different proposition from a pool of consultants rotating through your Slack channel, and it is the difference most often hidden behind a headline rate. Ask what they need from you as well: a recommendation engine build fails quietly when the client side has no engineer able to ship an event change inside a sprint.
Ask how they handle disagreement about scope. Personalisation scope grows by surface, and a fixed price only holds if the surface list is fixed, so the change process should be a written mechanism agreed at signature rather than a negotiation under deadline pressure. If you are still weighing whether to run this internally instead, our comparison of Eazyware and an in-house team sets out the trade honestly.
Questions not worth asking
How many data scientists do you have is not a useful question; team size predicts nothing about whether a ranking ships. Which model do you use is not useful either, because the correct answer changes every few months and any competent team benchmarks per task. Do you have experience in our industry is weak on its own: catalogue shape, traffic volume and data maturity predict the work far better than sector does.
Security questions are worth asking, but ask them properly rather than as a checkbox. The list in a security questionnaire for AI vendors is a better use of an hour than a generic certification question, particularly around where behavioural data is processed and how deletion requests propagate to trained artefacts.
When a vendor process is premature
If you have never personalised anything and your event data has never been audited, running a competitive process now will produce confident documents you cannot judge. Buy a small piece of work instead. A ten-day Sprint Zero at $3,250 or ₹2,00,000, credited to the build, returns the event audit and the test design, and a three-week ProofRun at $6,250 or ₹4,00,000 proves one surface. Then run the process with real numbers.
It is also premature if nobody owns the metric. Without an accountable owner, the best vendor in the market delivers a pilot that never rolls out, and the post-mortem blames the technology.
Related reading
What to put in a recommendation engine development RFP covers the document that precedes these conversations, the hidden costs of recommendation engine development lists the lines to interrogate in a quote, and our personalisation work with a D2C brand shows the sequence a good answer describes. If you would rather have the conversation directly, talk to us.
The vendor you want is the one who tells you what could go wrong before you have paid them anything.
Frequently asked questions
What is the single most revealing question to ask a recommendation engine vendor?
▾
Ask them to describe a recommendation engine they shipped, the surface it ran on, the primary metric, the holdout size and duration, and the result including the tests that came back flat. Vendors who have run a controlled test answer in specifics. Vendors who have only built demos answer with offline accuracy.
Should we insist on owning the trained model and pipelines?
▾
Yes. Insist on owning the code, the data pipelines, the trained model artefacts, the feature definitions, the evaluation sets and the documentation, held in your own repositories and cloud accounts. Anything less makes switching vendors expensive and leaves your behavioural data inside someone else's system.
How do you check a vendor's claims about improvement?
▾
Ask how the number was measured, not what it was. A credible claim names a randomised holdout, the percentage of traffic held back, the test window in full weeks and the guardrail metrics watched alongside. An improvement measured by comparing before and after periods is not evidence, because seasonality and campaigns are uncontrolled.