Vector search vs keyword search: which to choose and when
What is the difference between vector search and keyword search?
Keyword search matches the words in your query against the words in your documents; vector search matches meaning through embeddings. Keyword wins on identifiers and rare terms, vector wins on paraphrase and intent, and most production systems that work well run both and fuse the results.
Keyword search matches the words in your query against the words in your documents. Vector search matches meaning by comparing numerical embeddings, so passages that say the same thing in different words still rank. Keyword wins on part numbers, codes and rare terms; vector wins on paraphrase and intent. Most production systems run both.
This article sets out what each technique does mechanically, a side-by-side on nine dimensions, the query shapes where each clearly wins, what a retrieval layer costs to build and run at published Eazyware prices, and what you pay if you have to change your mind a year later.
What keyword search actually does
Keyword search, also called lexical or sparse retrieval, builds an inverted index: a map from every term in your corpus to the documents containing it. A query looks up each term, intersects the posting lists, and ranks the survivors with a scoring function. Across Lucene descendants and Postgres full-text search that function is a variant of BM25, which rewards documents where a query term is frequent and the term itself is rare across the corpus.
The consequence is that keyword search is literal and predictable. If a user types a part number, an invoice ID, an error code or a drug name, the engine finds documents containing that exact string and nothing else. It is cheap, well understood by operations teams, and explainable: you can point at the matched terms and say why a document ranked. What it cannot do is connect "my card is blocked" to a document titled "temporary freeze after failed PIN attempts". No shared terms means no match, and the user sees an empty result page.
What vector search actually does
Vector search converts text into an embedding, a list of numbers positioned so that passages with similar meaning sit close together. The query is embedded the same way and the engine returns its nearest neighbours by cosine similarity. Because nothing is matched literally, "card blocked" and "temporary freeze" land near each other with no word in common. The general technique is semantic search, and the store holding the vectors is a vector database.
The price of that flexibility is precision on exactly the cases lexical retrieval handles without effort. Embeddings compress meaning, and compression loses detail. A query for part number XR-4410B will often return XR-4411B just as confidently, because to the model those strings are nearly the same thing. Vector search also never returns nothing: it hands back the k nearest vectors however far away they sit, so a query with no good answer produces plausible-looking rubbish unless you set a similarity threshold.
Vector search vs keyword search: a side-by-side
| Dimension | Keyword search (BM25) | Vector search (embeddings) |
|---|---|---|
| Matching basis | Exact terms, stems and synonyms you configure | Semantic proximity in embedding space |
| Best query shape | IDs, codes, names, rare terms, Boolean filters | Questions, paraphrases, descriptions of intent |
| Recall on unseen phrasing | Poor unless synonym lists are maintained | Strong; this is the main reason to adopt it |
| Precision on identifiers | Exact | Weak; near-miss codes score almost identically |
| No good answer exists | Returns nothing, which is honest | Always returns k results, so a threshold is required |
| Index build cost | Fast, no model required | Embed every chunk, then re-embed on model change |
| Query cost | Milliseconds of CPU, no model call | One embedding call per query plus an ANN lookup |
| Explainability | Matched terms are visible to an auditor | A similarity score, which is hard to defend |
| Operational maturity | Decades of tooling and staff familiarity | Newer; index tuning is its own skill |
Where keyword search clearly wins
Keyword search wins wherever the user already knows the exact token they want. Catalogue lookup by SKU, ticket search by reference number, compliance corpora where a clause number must resolve exactly, and log search. It also wins when the corpus is small and well titled: a two-hundred-page policy handbook with honest headings does not need an embedding pipeline.
There is a quieter argument for lexical retrieval that rarely appears in vendor material. It degrades honestly. When it finds nothing it says so, the user rephrases, and nobody is misled. That is a better experience than a confident wrong answer, and in regulated settings it is also a better audit position.
Where vector search clearly wins
Vector search wins when the gap between how users write and how your content is written is wide. Support knowledge bases written by engineers and queried by customers, catalogues where shoppers describe a use case rather than a category, and anything feeding a language model. Retrieval-augmented generation is the largest single driver: a model needs passages that are topically right, not passages that share a token. How that pipeline behaves under real load is covered in why basic RAG fails in production.
It also wins on multilingual corpora. A customer typing in Hindi against English documentation gets nothing useful from BM25 and reasonable results from a multilingual embedding model. For Indian consumer products that difference is often the single finding that justifies the project.
The case for running both
For most production systems the honest answer is both, which the industry calls hybrid search. You run the lexical query and the vector query in parallel, normalise two incomparable score distributions, fuse the ranked lists, and rerank the top results. OpenSearch documents this pattern in its hybrid search guide, including the normalisation step that makes BM25 scores and similarity scores comparable at all.
Hybrid is not free. You maintain two indexes, two query paths and a fusion layer, and you need an evaluation set to prove the fusion beats either leg alone rather than merely feeling better in a demo. What it buys is the removal of the worst failure of each approach: the empty result page and the confident near miss. The longer treatment is in hybrid search: why vectors alone miss the answer.
How to choose: sort fifty real queries
Before you buy anything, take fifty real queries from your logs and sort them into buckets. The mix decides the architecture, and it decides it in an afternoon.
- Identifier queries. Order numbers, SKUs, case IDs, error codes. Keyword search, always. If this is more than a third of traffic, lexical is your primary leg and vectors are the assist.
- Natural language questions. "Can I return something after thirty days", "why was my claim rejected". Vector search, or hybrid with the vector leg weighted higher.
- Mixed queries. "Return policy for order 88421." These need both legs plus a metadata filter, and they are the reason pure vector deployments disappoint in retail and support.
- Filtered browse. "Blue running shoes under three thousand rupees." This is structured filtering, not search. Neither technique helps; your product metadata does.
- Queries with no good answer. Count these separately. If they are common you need a similarity threshold and an honest "we do not have this" reply more than you need a better index.
What does each cost to build and run?
Keyword search on a cluster you already operate is usually a few days of engineering and no new infrastructure. Vector search adds an embedding pipeline, a vector store, chunking decisions, permission filtering and an evaluation set, which is why we price it as a proper engagement. Retrieval and knowledge engineering at Eazyware starts at $14,000 or ₹8.8 lakh and runs to $49,000 or ₹32 lakh depending on corpus size, permission complexity and source systems. For search inside a product rather than behind an assistant, Eazy Search AI is the packaged version, and every starting figure sits on the pricing page.
Running costs differ in shape, not only in size. Keyword search consumes CPU and disk. Vector search consumes an embedding call per query, memory for the index, and a full re-embedding of the corpus every time you change embedding model. That last item is the line most budgets miss, and model deprecation schedules make it recurring. After launch, the AI add-on to a Care Plan at $750 or ₹40,000 a month covers re-indexing, retrieval evals and cost monitoring.
When vector search is the wrong choice
First, when your content is bad. Embeddings retrieve what exists. If the knowledge base is forty stale PDFs and a wiki nobody has edited since 2023, a vector index makes retrieval of stale content faster rather than better. Fixing the content is cheaper, and it is work you would have to do anyway.
Second, when the corpus is tiny. Under a few hundred short documents, a tuned BM25 index with a synonym list beats an embedding pipeline on accuracy and on total cost of ownership, and it ships in a week instead of a month.
Third, when explainability is a legal requirement. If an auditor can ask why document A ranked above document B, "the cosine similarity was 0.81" is a weak answer. Lexical matching leaves a trace a human can read, and in BFSI and healthcare that trace is sometimes the deciding factor.
What it costs to change your mind later
Reversal is asymmetric, and that should influence the first decision. Moving from keyword to hybrid is additive: the inverted index stays, you add an embedding pipeline and a fusion layer, and nothing already built is discarded. Budget four to eight weeks depending on corpus size.
Moving the other way is worse. Teams that went vector-first usually never built the metadata, synonym lists and field boosts that make lexical search good, so you are not restoring an index, you are building one from scratch while a live system underperforms. Keeping a golden question set from day one is what makes any of these moves measurable rather than argued.
A checklist before you commit
- Export a month of real queries and classify them into the six buckets above
- Count how many queries currently return nothing, and how many return the wrong thing confidently
- Decide whether permission filtering happens before or after retrieval, and write it down
- Pick an embedding model and note its deprecation policy before you index a million chunks
- Build a golden question set of real queries with the passages that should win
- Name the owner of the monthly relevance review, because relevance drifts with your catalogue
Related reading
RAG vs fine-tuning answers the question most teams ask after they settle retrieval, pgvector vs Pinecone vs Qdrant covers where the vectors live, and GraphRAG explained covers cases where neither is enough. The rest of our side-by-sides sit on the comparison hub.
Choose your retrieval layer from your query log rather than from a vendor deck, because the shape of what people actually type decides this, and for most teams it decides in favour of both.
Frequently asked questions
Is vector search better than keyword search?
▾
Neither is universally better. Vector search recovers meaning when the user's words differ from the document's words. Keyword search is exact on identifiers, codes and rare terms, and it tells you honestly when there is no match. Measured against a real query log, most corpora need both legs fused into one ranked list.
Do I need a dedicated vector database to run vector search?
▾
Not always. Postgres with the pgvector extension handles millions of vectors comfortably and keeps embeddings next to your relational data and its access rules. A dedicated vector database earns its place when you need very large indexes, heavy filtering at query time, or managed scaling. Start with the system your team already operates.
How do I know whether my search is actually improving?
▾
Build a golden question set of fifty to two hundred real queries paired with the passages that should be returned, then measure recall at k, precision and groundedness before and after every change. Without that set you are judging search by anecdote, and every change feels like an improvement to whoever made it.