Get Started
Home / Solutions / Product Matching
Catalog identity and comparability

Product Matching That Makes Every Comparison Trustworthy

Know which products are identical, which are comparable, and which should never be compared. Kvetoiq combines identifiers, normalized attributes, text, images, business rules, confidence scoring, and human review.

Explainable decisions Typed relationships Human review controls
Match Confidence Workbench
Reason codes visible
Approved catalog Wireless Earbuds X2
Model
WX2-2026
Color
White
GTIN
012345678901
Pack
1 unit
Retailer listing X2 Wireless Buds
Model
WX2
Color
White
GTIN
012345678901
Quantity
Single
Relationship Exact Product
GTIN exact Model normalized Color exact Pack exact
Approved downstream uses Price, availability, and digital shelf comparison
Exact Verified product identity
Comparable Clearly labeled relationship
Explainable Signals and conflicts retained
Reviewable Ambiguity routed to people
Maintainable Versioned as catalogs change
What is product matching?

Product matching identifies and classifies relationships between product records across catalogs, retailers, marketplaces, sellers, or internal systems.

A reliable output preserves whether the relationship is exact, variant, comparable, substitute, private label, duplicate, or unmatched.
The data foundation

The match determines whether every downstream comparison can be trusted.

Price, availability, content, seller, assortment, and MAP analysis become unreliable when the wrong products are compared. Matching is not a setup detail. It is the identity layer beneath reliable retail data.

The cost of a wrong match

Match the wrong products and every decision after that is wrong.

Catalogs rarely share clean identifiers, naming conventions, units, taxonomies, or pack structures. A plausible-looking pair can still represent a different model, quantity, formulation, market, or product generation.

01
False price comparisons Different products or pack sizes create misleading gaps.
02
Incorrect MAP alerts An offer is evaluated against the wrong governed item.
03
Broken assortment analysis Exact overlap, variants, substitutes, and unique items are confused.
04
Contaminated digital shelf metrics Content, availability, and reviews attach to the wrong identity.
Relationship taxonomy

Not every similar product is the same product.

Kvetoiq keeps each relationship type explicit so the result can be approved only for suitable downstream uses.

01

Identifier-Exact Match

The same verified UPC, GTIN, EAN, ASIN, MPN, or manufacturer identifier.

Exact price, availability, and listing comparison
02

Inferred Exact Match

The same product established through consistent brand, model, attributes, text, imagery, and specifications.

Exact comparison after required validation
03

Variant Match

The same product family with a different size, color, flavor, capacity, or configuration.

Family and variant analysis
04

Pack and Quantity Match

Equivalent units represented through different pack structures, quantities, weights, or volumes.

Normalized unit and pack comparison
05

Comparable Product

Different products that satisfy an agreed category-specific comparison framework.

Competitive and assortment benchmarking
06

Functional Substitute

Different products that may address a similar need without being directly equivalent.

Substitution and category analysis
07

Private-Label Equivalent

A retailer-owned product mapped to a branded comparable using category-specific attributes.

Private-label competitive analysis
08

Duplicate Record

Two or more records representing the same underlying product inside one data environment.

Deduplication and catalog quality
Matching signals

Multiple signals, one explainable decision.

The useful signal mix changes by category, source quality, identifier availability, and match relationship.

ID

Identifiers

Use deterministic product and manufacturer codes when available and trustworthy.

  • UPC, GTIN, and EAN
  • ASIN and marketplace ID
  • MPN and model number
  • Internal SKU
BRAND

Brand and Manufacturer

Normalize ownership, product-line, and private-label relationships.

  • Brand and sub-brand
  • Manufacturer
  • Product line
  • Private-label owner
NLP

Text and Semantics

Compare normalized meaning rather than relying on raw character similarity.

  • Title and description
  • Bullet points
  • Abbreviations
  • Multilingual equivalents
ATTR

Attributes and Specs

Evaluate the structured characteristics that define equivalence by category.

  • Size and color
  • Material and capacity
  • Technical specifications
  • Formula and generation
PACK

Pack and Unit Data

Normalize how quantity, weight, volume, and multipack structures are represented.

  • Pack count
  • Unit quantity
  • Weight and volume
  • Per-unit conversion
IMAGE

Visual Evidence

Use imagery as supporting evidence while accounting for reused or outdated assets.

  • Primary product image
  • Packaging
  • Visual features
  • Image consistency
TAX

Taxonomy and Category

Keep candidate generation and comparison inside the correct product context.

  • Product type
  • Retail category
  • Classification
  • Category rules
CHECK

Contextual Checks

Use contextual fields to detect conflicts, not as the sole basis for identity.

  • Price range
  • Seller and market
  • Currency and geography
  • Availability context
Key features

Everything required for reliable cross-catalog matching.

Prepare the data, establish relationships, control uncertainty, and maintain the match as catalogs change.

01Prepare 02Match 03Control 04Maintain
PROFILE
Prepare

Catalog Profiling

Assess schema quality, identifiers, missing fields, duplication, category coverage, and source differences.

  • Field completeness
  • Identifier coverage
  • Duplicate patterns
  • Category distribution
NORM
Prepare

Data Normalization

Standardize product text, units, brands, pack quantities, values, and formats before comparison.

  • Brand normalization
  • Unit conversion
  • Pack parsing
  • Text standardization
TAX
Prepare

Taxonomy Alignment

Map source categories into a controlled product framework for relevant candidate generation.

  • Category mapping
  • Product-type rules
  • Classification standards
  • Category-specific attributes
ID
Match

Identifier Matching

Resolve reliable universal, manufacturer, marketplace, and internal product identifiers.

  • Exact identifier joins
  • Identifier validation
  • Conflicting ID checks
  • Source hierarchy
NLP
Match

Attribute and NLP Matching

Combine semantic text comparison with normalized category-defining attributes.

  • Title and description
  • Attribute similarity
  • Specification conflicts
  • Multilingual matching
VISION
Match

Image-Assisted Matching

Use product and packaging imagery as an additional signal when text or identifiers are incomplete.

  • Visual similarity
  • Packaging recognition
  • Asset conflict detection
  • Image quality checks
TYPE
Control

Match-Type Classification

Label the precise relationship instead of reducing every valid pair to a generic match.

  • Exact and inferred exact
  • Variant and pack
  • Comparable and substitute
  • Private label and duplicate
CONF
Control

Confidence Scoring

Combine supporting signals, conflicts, source quality, and category rules into calibrated bands.

  • Signal weights
  • Conflict penalties
  • Threshold controls
  • Use-case permissions
REVIEW
Control

Human Review Queue

Route ambiguous, high-impact, and conflicting pairs to trained reviewers with full context.

  • Prioritized queue
  • Side-by-side evidence
  • Reviewer decision
  • Reason capture
WHY
Maintain

Explainable Match Reasons

Retain which signals supported the decision and which fields created uncertainty.

  • Reason codes
  • Supporting fields
  • Conflicting fields
  • Decision trace
VERSION
Maintain

Versioning and Monitoring

Revalidate relationships as products, listings, attributes, rules, and catalogs change.

  • Ruleset version
  • Change detection
  • Match history
  • Revalidation schedule
DATA
Maintain

Flexible Data Delivery

Send match relationships, confidence, reasons, and review status into downstream systems.

  • CSV, Excel, and JSON
  • API and webhooks
  • Cloud and warehouse
  • Database delivery
Matching workflow

From messy catalogs to maintained product relationships.

Each stage improves reliability and preserves the evidence behind the final decision.

01

Profile

Inspect schemas, fields, identifiers, and quality.

02

Normalize

Standardize brands, units, packs, text, and taxonomy.

03

Generate

Create plausible candidate pairs efficiently.

04

Score

Evaluate identifiers, attributes, text, images, and rules.

05

Classify

Assign the correct relationship type.

06

Control

Apply confidence bands and downstream permissions.

07

Review

Resolve ambiguous and high-impact pairs.

08

Maintain

Deliver, version, monitor, and revalidate.

Evaluation framework

Accuracy is not one number.

A trustworthy evaluation separates different error types, relationship classes, categories, confidence bands, and business consequences.

Precision

How many declared matches are correct?

High precision limits false comparisons and is critical for price and MAP workflows.

Correct matches ÷ All declared matches
Recall

How many valid matches were found?

Higher recall helps discover more duplicates, overlap, and assortment relationships.

Found valid matches ÷ All valid matches
False Positives

How often are unrelated products joined?

False positives can contaminate pricing, compliance, availability, and performance analysis.

Incorrectly accepted product pairs
False Negatives

How often is a valid relationship missed?

False negatives reduce catalog coverage and hide relevant competitive relationships.

Valid pairs incorrectly rejected or missed
Coverage

How much of the catalog received a usable decision?

Coverage distinguishes matched, reviewed, rejected, and unresolved records.

Records with usable status ÷ Total records
Review Rate

How much requires human validation?

Review volume reveals where source quality or category complexity limits automation.

Review-required pairs ÷ Candidate pairs
Confidence and control

Let confidence determine the next action, not hide the uncertainty.

A score should control whether a relationship is accepted, reviewed, rejected, or withheld from a specific downstream workflow.

Match-type-specific thresholds
Category-specific signals
Conflict-aware scoring
Use-case permissions
Human review escalation
Ruleset and model versioning
Verified

Strong deterministic or reviewed evidence. Approved for defined downstream use.

High Confidence

Multiple consistent signals with no material conflicts above the acceptance threshold.

Review Required

A plausible relationship with missing, ambiguous, or conflicting evidence.

Rejected

Material conflicts indicate that the pair should not be treated as valid.

Unmatched

No candidate reached the required threshold for the requested relationship.

Pack normalization example Comparable unit quantity
One 24-Pack 24 units × 10 g
=
Two 12-Packs 2 × 12 units × 10 g
Total units 24 = 24
Total quantity 240 g = 240 g
Relationship Pack equivalent
Pack and variant logic

Normalize the quantity before comparing the offer.

A title can describe “24 count,” “2 × 12,” or “24 individual units” while representing the same total quantity. Kvetoiq parses pack structure, unit count, weight, volume, size, and variant attributes before assigning the relationship.

Business use cases

One identity layer supports every retail data workflow.

PRICE

Competitor Price Monitoring

Compare prices only after equivalent products, variants, and pack quantities are confirmed.

Explore Price Monitoring →
MAP

MAP Violation Detection

Connect observed offers to the correct governed product before evaluating policy thresholds.

Explore MAP Monitoring →
SHELF

Digital Shelf Analytics

Create one cross-retailer identity for content, availability, seller, rating, and visibility analysis.

Explore Digital Shelf Analytics →
BRAND

Brand Protection

Connect suspicious listings and seller offers to the correct approved product identity.

Explore Brand Protection →
GAP

Assortment Gap Analysis

Separate exact overlap, variants, substitutes, private labels, and unique products.

CLEAN

Catalog Deduplication

Find duplicate and near-duplicate records inside product systems and source catalogs.

ONBOARD

Marketplace Onboarding

Connect seller offers with existing catalog identities or route genuinely new products.

LABEL

Private-Label Analysis

Map retailer-owned products to branded comparables without claiming exact equivalence.

ENRICH

Product Data Enrichment

Use matched sources to identify missing attributes and catalog inconsistencies.

SEARCH

Product Discovery Analysis

Connect retailer search visibility to the correct products, variants, and product families.

Explore Share of Search →
Output schema

Deliver the relationship, evidence, conflicts, and permitted use.

A match output should be auditable and useful downstream, not merely two IDs and an unexplained score.

Output field Purpose
Source and target IDs Connect the original product records without losing source identity.
Relationship type Exact, variant, pack, comparable, substitute, private label, duplicate, rejected, or unmatched.
Confidence and band Provide the calibrated score and operational decision state.
Reason and conflict codes Show which fields support the decision and which fields disagree.
Normalized product fields Preserve brand, model, attributes, pack quantity, taxonomy, and other standardized values.
Review status Record reviewer decision, notes, and resolution when human validation is required.
Version and timestamps Track ruleset, creation, validation, and latest recheck context.
Allowed downstream uses State whether the pair can support price, MAP, digital shelf, assortment, or other analysis.
CSV Excel JSON REST API Webhook Amazon S3 Google Cloud Azure Snowflake BigQuery SQL Database
Catalog and marketplace coverage

Match products across the sources relevant to your market.

AMZ

Amazon Catalogs

Connect ASINs, offers, variants, packs, brands, and product attributes with external records.

Explore Amazon data →
WMT

Walmart Catalogs

Map retailer and marketplace products using identifiers, titles, attributes, images, and sellers.

Explore Walmart data →
SHOP

Direct Stores

Compare marketplace listings with brand, retailer, distributor, and competitor store catalogs.

Explore Shopify store data →
TTS

TikTok Shop Catalogs

Match shop products, variants, sellers, and offers with approved catalog records.

Explore TikTok Shop data →
FK

Flipkart Catalogs

Connect product listings, variants, sellers, attributes, and offers across catalogs.

Explore Flipkart data →
EBAY

eBay Catalogs

Resolve marketplace listings, conditions, sellers, variants, and model identities.

Explore eBay data →
ETSY

Etsy Catalogs

Classify handmade, customized, vintage, and comparable product relationships.

Explore Etsy data →
BBY

Best Buy Catalogs

Match electronics using model numbers, specifications, variants, and identifiers.

Explore Best Buy data →
WAY

Wayfair Catalogs

Connect furniture and home products using dimensions, materials, styles, and variants.

Explore Wayfair data →
ABA

Alibaba Catalogs

Normalize supplier products, specifications, pack structures, and minimum quantities.

Explore Alibaba data →
AEX

AliExpress Catalogs

Map product variants, sellers, specifications, offers, and cross-border listings.

Explore AliExpress data →
WISH

Wish Catalogs

Match listings through normalized titles, attributes, images, sellers, and variants.

Explore Wish data →
NE

Newegg Catalogs

Resolve electronics and component identities using models, specifications, and sellers.

Explore Newegg data →
WEB

Category Retailers

Build tailored matching schemas for grocery, beauty, electronics, fashion, pharmacy, and other categories.

Explore Custom Extraction →
SCALE

Large Source Networks

Coordinate broader catalog collection, normalization, matching, and maintenance programs.

Explore Enterprise Crawling →
INTERNAL

Internal Product Systems

Match external records with PIM, ERP, data warehouse, seller, and approved catalog identities.

Responsible AI and review

Automate the clear decisions. Expose the uncertain ones.

Kvetoiq combines deterministic rules, machine-assisted signals, confidence thresholds, validation data, and human review. Match quality is evaluated by relationship type and use case rather than hidden behind one universal accuracy claim.

Relationship types preserved
Supporting and conflicting signals visible
Ambiguous pairs routed to review
Rules and decisions versioned
Frequently asked questions

What teams ask about Product Matching.

What is product matching?

Product matching identifies and classifies relationships between product records across catalogs, retailers, marketplaces, sellers, or internal systems. Relationships may be exact, variant, pack equivalent, comparable, substitute, private label, duplicate, rejected, or unmatched.

How does product matching work?

Catalogs are profiled and normalized, plausible candidate pairs are generated, identifiers and product signals are scored, a relationship type is assigned, confidence controls are applied, ambiguous pairs are reviewed, and approved results are delivered and maintained.

What is an example of product matching?

A retailer listing titled “X2 Wireless Buds, White” may be matched with an approved “Wireless Earbuds X2” record when the GTIN, brand, normalized model, color, specifications, and pack quantity agree.

What is the difference between exact and similar product matching?

An exact match represents the same underlying product. Similar matching covers defined relationships such as variants, comparable products, substitutes, or private-label equivalents. Similar products should not automatically be used for exact price comparison.

Can products be matched without UPCs or GTINs?

Yes. Inferred matching can combine brand, manufacturer, model, title, description, attributes, specifications, pack data, taxonomy, imagery, and contextual checks. Missing identifiers usually increase uncertainty and review requirements.

Which product attributes are used?

The fields depend on the category and may include model, size, color, flavor, capacity, material, dimensions, technical specifications, formula, generation, pack count, weight, volume, and other category-defining properties.

How are images used in product matching?

Images can provide supporting evidence through product, packaging, shape, label, and visual-feature similarity. They should be combined with other signals because images may be reused, outdated, incomplete, or visually similar across different products.

How are pack sizes and multipacks matched?

Pack count, unit quantity, weight, volume, and multipack structure are parsed and normalized. One 24-pack may be identified as pack-equivalent to two 12-packs while preserving that the listed pack structures differ.

How are product variants handled?

Products from the same family can be linked while retaining variant fields such as size, color, flavor, capacity, model generation, or configuration. Variant relationships remain separate from exact-SKU relationships.

Can private-label products be matched?

Yes, when a category-specific comparability framework is defined. Private-label products should be labeled as comparable or equivalent for the intended analysis, not misrepresented as identical branded products.

What is a product matching confidence score?

A confidence score summarizes supporting signals, conflicts, source quality, category rules, and model or ruleset outputs. It should map to an operational band such as verified, high confidence, review required, rejected, or unmatched.

How is product matching accuracy measured?

Evaluation should include precision, recall, false-positive rate, false-negative rate, coverage, review rate, and confidence calibration. Results should be segmented by relationship type, category, source, identifier availability, and confidence band.

What happens when a match is uncertain?

The pair can be routed to a review queue with side-by-side records, supporting signals, conflicting fields, suggested relationship type, and downstream impact. It should not be silently accepted.

Is human review available?

Yes. Human review can support ambiguous, conflicting, high-value, private-label, pack-complexity, and category-specific cases. Review scope is designed around risk and required precision.

Can catalogs in different languages be matched?

Multilingual matching can use translated or normalized text, identifiers, brands, models, attributes, specifications, taxonomy, images, and market-specific rules. Feasibility depends on source quality and category complexity.

Can Product Matching support price monitoring?

Yes. Product Matching establishes which products and pack quantities are valid for exact or comparable price analysis. Match relationship, confidence, and comparison rules should remain attached to the price record.

How is match data delivered?

Delivery can include CSV, Excel, JSON, APIs, webhooks, databases, cloud storage, warehouses, BI platforms, or integration with approved internal product systems.

How do we begin a Product Matching project?

Share the source and target catalogs, category, record counts, identifiers, desired relationship types, downstream use case, required precision, review expectations, and preferred output format. Kvetoiq will profile the data before defining the production workflow.

Test the match quality

See the relationship, evidence, conflicts, and confidence before you scale.

Share sample source and target catalogs, your downstream use case, and the relationship types you need. Kvetoiq will profile the data and prepare a reviewable Product Match sample.

LET'S TALK

Tell us what market decision you need to make next.

Share the platforms, categories, competitors, SKUs, regions, or business questions you care about. KVETOiQ will help define the right data strategy, output format, and operating cadence.

  • Pricing and promotion monitoring
  • Marketplace and seller intelligence
  • Digital shelf and search visibility
  • Review sentiment and customer intelligence

    Get Your Custom Data

    No spam
    Response within 24 hrs