When a search does not produce a good result, the search technology is often the first to be blamed. However, a large part of the problems arise before: in product titles, categories, attributes, variants, identifiers and availability. Search can normalize words and calculate relevance, but cannot make missing product knowledge reliable. Those who want better results structurally must therefore improve search data and product data together.

What product data does to search quality

A search engine compares the customer’s need with information about the assortment. The more complete and consistent that information is, the better products can be found, filtered and arranged. For example, a product with only a commercial title is difficult to match reliably on material, application or technical feature.

Product data also determines the context. The same term can mean something different in different categories. A size with clothing works differently than a size with furniture; ‘power’ is relevant in electronics but not with care products. Categories and attributes help search understand those differences.

This does not mean that every product feed must be perfect before search can deliver value. However, it must be clear which data fields are essential for the most important customer questions. Start there, instead of cleaning up all the fields at once.

Problem 1: product titles are incomplete or inconsistent

A good product title contains sufficient distinctive information without becoming a list of each available feature. When make, model, product type, size and execution are in random order, matching becomes less predictable and results are harder for customers to compare.

Also pay attention to internal abbreviations and supplier names. A warehouse code can be operationally useful, but a customer doesn't say anything. Conversely, a commonly used customer term may be missing because the supplier uses a technical name. Where possible, store both the source value and a customer-centric view so that search can connect to both.

Create title rules by product group. A usable pattern for a laptop differs from that for a dress or a spare part. Then check that the parts remain available as separate fields; only a composite title makes filtering and targeted boosting more difficult.

Example: the same information, different usability

Model X black 42 brand Y may contain all the words, but requires a lot of interpretation. Brand Y Model X running shoe black – size 42’ makes product type and distinguishable. The example is not a universal title recipe; it shows why consistent meaning is more important than adding as many terms as possible.

Problem 2: Attributes missing, bumping or using free text

Attributes are needed for filters and for search queries in which customers combine multiple properties. If material is missing from half of the products, a material filter can exclude relevant items. When values such as ‘dark blue’, ‘navy’ and ‘navy’ are next to each other without normalization, fragmented filters arise.

Free text is flexible, but difficult to process consistently. For values that customers filter or compare, controlled fields are usually more useful. At the same time, keep sufficient original detail information; too aggressive normalization can lose meaning. For example, a water resistance class is not the same as a general value ‘water resistant’.

Check by category which attributes are mandatory, recommended or irrelevant. Make data quality visible as coverage: what proportion of saleable products has a valid value? Combine that coverage with search volume to determine priorities.

  • Missing values for frequently viewed or frequently sought after products.
  • Multiple writing modes for the same meaning.
  • One field in which several units are mixed.
  • Values that are internally correct but incomprehensible to customers.
  • Attributes associated with the wrong product group.

Problem 3: Categories follow the organization instead of the customer

An internal or logistics category tree is not automatically suitable for product discovery. Suppliers, warehouses and financial systems group items for reasons other than customers. When search uses that structure without translation, results and filters can feel illogical.

A customer-centric taxonomy helps with query understanding, filter selection and ranking. That doesn't have to mean that all source systems have to use the same tree. A translation layer can link internal categories to sales categories, use cases and alternative inputs.

Pay attention to products that fit in multiple contexts. A water bottle can be relevant under sports, travel and office. One mandatory category should not unnecessarily limit that findability. Work where appropriate with main category, additional classifications, and attributes describing the usage context.

Problem 4: variants, identifiers and availability are not reliable

Variants should be presented as one logical product without hiding important differences. If each color and size appears as a separate result, a result page quickly becomes filled with almost identical articles. If variants are too strongly combined, a search for a specific size or version can wrongly show a product that is not available.

SKUs, EANs and model codes are important for customers who know exactly what they are looking for, especially in B2B and parts. Normalize spaces, dashes, and capitals for matching, but save the original value for display and control. Protect exact identifier matches from broad synonym or spelling rules.

Price and stock must be updated in a timely manner. A highly ranked but not available product can sometimes still be useful, for example with an alternative or expected delivery date. The choice must be conscious. Outdated status that inadvertently affects ranking and filters is a data problem.

Check parent and variant level separately

Record what information belongs at the product family level and which differs per variant. Title and brand can be shared, while size, color, stock and sometimes price are variant specific. Search and storefront need to understand the same model to avoid conflicting results.

Use search data to improve product data in a targeted manner

A complete data audit can be bulky. Search data helps determine the order. Start with commonly used queries, key categories, unsuccessful searches, and terms with remarkably low click ratio. Then investigate which missing or inconsistent data causes multiple problems at once.

Make each finding a concrete data rule. Do not ‘improve titles’, but for example ‘for all saleable running shoes, brand, model, product type and target group are filled separately’. Link the rule to an owner and validate in source, feed, index, and storefront. A correction in one intermediate layer that disappears at the next import is not a sustainable solution.

After a change, do not only measure whether more products are found. Also check relevance, filtering behavior, product clicks and conversion. More results are not automatically better results. The goal is to make the right products more accessible.

  • Prioritize customer impact and number of affected products.
  • Solve problems as close to the reliable source as possible.
  • Add automatic validations for required fields and values.
  • Test important queries again after each major feed or taxonomy change.
  • Establish exceptions so that temporary repairs do not become permanent fault.

Search quality starts with useful product knowledge

A search engine can do a lot: normalize spelling, expand terms, arrange results and take behavior with them. However, it cannot reliably determine that a product has a property when that information is not available anywhere. Therefore, structural product data improvement not only makes one query better, but often a whole group of search queries, filters and categories.

Findoviq can make visible where customers get stuck and what queries deserve priority. Sustainable improvement is achieved when search specialists, e-commerce, product management and source data owners use the same signals and solve problems at the right source.

Frequently asked questions

Which product data is minimally needed for good online store search?

This varies by assortment. Usually, a clear product title, brand, product type, category, sales status, price and relevant distinctive attributes are needed. Model codes and specifications are often essential for technical products; size, color, material and target group play a larger role for fashion.

Can AI automatically complete missing product data?

AI Can make proposals or structure information from existing texts, but such outcomes must be verifiable. Properties that are not reliably recorded anywhere should not be invented as fact. Therefore, use automation with validation rules, source reference, and human control where the risk requires it.

Should product data be restored in the PIM or in the search engine?

Preferably resolve an error in the most reliable source systems, so that all customers benefit and the correction does not disappear with a subsequent import. Search-specific standardization can be useful, but should not become an invisible replacement for structural source quality.

How do you measure whether better product data improves search?

Use a fixed query set and compare zero-result searches, relevance, filter availability, product clicks, and possibly reliable linked conversion. Also check that the change does not cause unwanted results in other categories.

Further reading

View product data and indexing →Read how to catch typos in a controlled manner →

More practical insights

View all FINDOVIQ items →