azyware
Technology

Search that understands intent: semantic search for catalogues

EZ
Eazyware
· 7 min read
Quick answer

What should you know about ecommerce semantic search?

Semantic search interprets a shopper's intent and matches products by meaning, with hybrid retrieval for exact matches. "Something warm for a Bengaluru winter" finds light jackets; "SKU 4471" still finds SKU 4471. The build combines embeddings, keyword search, a re-ranker and filters, scored on real queries.

Ecommerce semantic search is site search that understands what a shopper means rather than which words they typed. Keyword search returns nothing for "gift for a friend who runs" and everything for "shirt"; semantic search reads the first as a category, a price range and an occasion, and narrows the second by what the shopper does next. This article explains how AI product search works, why hybrid retrieval is not optional, how to handle the catalogue and its attributes, how to measure improvement, and what a build involves.

What ecommerce semantic search is and why it matters

Site search is the highest-intent surface a store has: a visitor who types is telling you what they want. Most catalogue search still runs on keyword matching with manual synonyms, so it fails on natural phrasing, misspellings, regional vocabulary, and any query that describes a need rather than a product. Semantic search embeds queries and products into the same vector space so "running shoes for flat feet" lands near products described as stability trainers even if the phrase never appears. The improvement shows up as fewer zero-result searches, more clicks on the first results and higher conversion from search sessions. The retail industry page lists search among the retail systems we build.

Keyword, semantic and hybrid search compared

QueryKeyword onlySemantic onlyHybrid with re-ranking
SKU 4471 or a brand model numberExact hitMay miss; numbers embed poorlyExact hit, boosted
"warm jacket under 3000"Matches "warm" and "jacket", ignores priceFinds jackets and knits, ignores priceFinds the right products, applies price as a filter
Misspelling or transliteration ("kurtaa", "chappal")Zero results without synonymsUsually correctCorrect, with keyword fallback
"something for a housewarming"Zero resultsFinds gift-appropriate home itemsFinds them and ranks by popularity and margin rules
Very short query ("shirt")Thousands of results, arbitrary orderThousands, arbitrary orderRanked by session context, popularity and business rules

Why hybrid retrieval is not optional

Embeddings capture meaning and lose specifics. A model number, a brand name in an unusual spelling, a colour code or a size all match poorly by vector similarity and perfectly by keyword. Shoppers use both kinds of query, often in the same session, so the search system runs both retrievers, merges the candidate sets, and hands them to a re-ranker that scores each product against the actual query. Hybrid retrieval is the single largest improvement over either method alone, and it is cheap. The general case is covered in hybrid search: why vectors alone miss the answer; the retail specifics are attribute filters and business rules on top.

The catalogue is the model's input

Semantic search is only as good as the text and attributes it embeds. Product titles written for a marketplace feed, missing attributes, and descriptions that say nothing about use or fit all degrade results. Before embedding, the catalogue is enriched: attributes normalised (sizes, colours, materials in one vocabulary), missing attributes inferred from images or descriptions with a human check on a sample, and a short use-oriented description generated where none exists and reviewed. Multilingual catalogues, or shoppers who search in Hindi or Tamil for products listed in English, need an embedding model tested on that combination rather than assumed to work.

Intent interpretation: filters, not just vectors

Much of a query is structured intent. "Under 3000" is a price filter. "For kids" is a category. "Blue" is an attribute. "Delivery by Saturday" is a fulfilment constraint. A query-understanding step extracts these into filters and leaves the remainder for semantic matching, so "warm jacket under 3000" becomes a vector search for warm outerwear with a price filter applied at retrieval, not afterwards. A small language model or a fine-tuned classifier does this in tens of milliseconds; the extraction is tested on real queries and its errors are visible in the evaluation set.

Ranking: relevance meets the business

Once candidates are retrieved and re-ranked for relevance, business rules apply: in-stock first, margin and campaign boosts within bounds, personalisation for recognised shoppers, and diversity so the first page is not ten variants of one product. These rules belong to the merchandising team and should be editable without a deployment. Session context helps short queries: a shopper who has been browsing women's footwear and types "black" almost certainly means black women's footwear. The session ranking model is described in real-time ranking: personalising within a single session.

Measuring site search improvement

A golden set of a few hundred real queries from search logs, each with the products a merchandiser judges relevant, is the test bench. Before any change, the current search is scored on it; every retrieval or ranking change is scored again. Offline metrics (whether relevant products appear in the top results, and how high) are joined by online ones: zero-result rate, click-through on the first page, add-to-cart from search, and conversion of search sessions, measured in an A/B test against the existing search. The measurement discipline is the same as for any retrieval system, described in how to measure RAG quality.

Infrastructure: what runs where

The search engine holds the keyword index, the vector index and the attribute filters together. OpenSearch and similar engines support all three in one system, which keeps latency low and operations simple; a separate vector database is justified only at very large scale. Embeddings are computed at catalogue update and cached; query embeddings are computed per request. The re-ranker runs on a small model that scores a few dozen candidates in well under the latency budget. Everything degrades gracefully: if the vector index is unavailable, keyword search still serves.

Conversational search and where it stops

Once query understanding and hybrid retrieval exist, a conversational layer is a small addition: the shopper asks "do you have this in a bigger size" or "what goes with it" and the system rewrites the follow-up into a full query with the context of the previous results. This is useful on mobile and in WhatsApp, where typing a complete query is a chore. It stops being useful when the conversation becomes a chatbot that talks instead of showing products. Every turn should end in a result set, a filter change or a single clarifying question; a search assistant that produces paragraphs has lost the shopper. The same layer can answer product questions from the catalogue's own data, sizes, materials, care instructions, with the answer grounded in the listing rather than generated.

A worked example

A home and lifestyle retailer with a large catalogue had keyword search with a synonym list maintained by one person. Zero-result searches were common and the search logs were full of natural phrasing and regional words the synonym list had never caught. The build started by enriching the catalogue: attributes normalised, use-oriented descriptions generated and reviewed for the top categories. Hybrid retrieval with a re-ranker replaced the keyword-only engine, with query understanding extracting price, category and colour into filters. A golden set from the logs scored the old and new systems before launch; an A/B test ran on live traffic afterwards. Zero-result searches fell sharply, the merchandising team gained a rules interface instead of a synonym file, and the retailer now reviews the golden set quarterly as the catalogue changes.

Team and timeline

Catalogue search AI is a retrieval and knowledge engineering build: a search engineer for retrieval and ranking, a data engineer for catalogue enrichment and indexing, and a front-end engineer for the search interface and analytics, over six to ten weeks. The first two weeks build the golden set and enrich the catalogue; the middle weeks tune hybrid retrieval and query understanding against the set; the last weeks add business rules, the merchandising interface and the A/B test. Starting price is $14,000 / ₹8.8L; personalised ranking adds a personalisation engine scope. The Essential Care Plan covers re-indexing, model updates and golden-set reviews. Prices are on the pricing page, and a three-week ProofRun that scores hybrid search against your current engine on your own queries is the low-risk way to start.

Before you start: a checklist

  • Export search logs and count zero-result queries and their phrasing
  • Audit catalogue attributes for completeness and vocabulary consistency
  • Decide the languages and scripts shoppers actually search in
  • Build a golden set of real queries with merchandiser-judged relevant products
  • List the business rules ranking must respect and who owns them
  • Confirm the latency budget the storefront can tolerate
  • Plan the A/B test against the current search before launch
  • Agree the fallback behaviour if any component is unavailable

Glossary

  • Embedding: a numeric representation of text such that similar meanings are close together
  • Hybrid retrieval: running keyword and vector search together and merging the candidates
  • Re-ranker: a model that scores retrieved candidates against the query for final ordering
  • Query understanding: extracting filters and intent from a query before retrieval
  • Golden set: real queries with judged relevant products, used to score every change
  • Zero-result rate: the share of searches that return nothing

Why basic RAG fails in production covers the same retrieval engineering for documents; Shopify personalisation without a platform subscription covers the ranking models search can share; pgvector vs Pinecone vs Qdrant covers the vector store decision.

Embed a catalogue worth embedding, run keyword and vector search together, and score every change on real queries; the shopper who typed is the one most ready to buy.

Frequently asked questions

Does semantic search replace keyword search?

▾

No. It runs alongside keyword search in a hybrid system, because model numbers, brand names and codes match better by keyword while natural phrasing matches better by meaning. A re-ranker orders the merged results.

How is ecommerce search improvement measured?

▾

Offline on a golden set of real queries with judged relevant products, and online by zero-result rate, first-page click-through, add-to-cart and conversion from search in an A/B test against the current engine.

Can semantic search handle Hindi or Tamil queries on an English catalogue?

▾

Often, with a multilingual embedding model tested on your catalogue and query mix; the golden set must include those queries so the result is measured rather than assumed.