Search Evaluation: proof that relevance is really getting better.
A few good example queries are insufficient. Build a fixed evaluation set and combine offline relevance with actual customer behavior.
Representative query set
Include top queries, long-tail, zero-result searches, brands, SKUs, categories, typos and natural questions.
Relevance judgments
Determine per query which results are relevant, partially relevant or unsuitable. This way you can compare changes reproducibly.
Regression tests
Check that improvements for one query group do not damage important exact matches or other segments.
Online metrics
CTR, Reformulations, add-to-cart and conversion give additional proof, but are also affected by price, stock and merchandising.
Segmentation
Compare devices, markets, stores, and query types separately when behavior differs structurally.
Decision-making
Promote a change only when pre-selected success criteria are sufficiently met and guardrails remain intact.