Skip to main content

Real Data Flow (RDF)

Real Data Flow (RDF) builds catalogue data out of the shelf images you have already sent to ZIA. Instead of asking you to photograph products in a studio, it groups together visually similar products found across your historical and production images, and proposes those groups as catalogue data for review.

Because image recognition treats your products and your competitors' products identically, RDF is the cheapest route to competitor coverage: the products are already in the images you have been sending β€” they just aren't in your catalogue yet.

What it is used for​

RDF runs in two distinct flows, and it is worth being clear about which one you want.

Increase coverage β€” find new items​

Finds products appearing in your images that are not in your catalogue, and proposes them as new catalogue items.

Use this for:

  • Onboarding competitor products for share-of-shelf and adjacency analysis
  • Catching products that entered the range without being added to the catalogue
  • Populating a thin catalogue from real shelf data

Where possible, proposed new items arrive with suggested metadata already filled in β€” barcode, name and size β€” so a new item is a review-and-confirm step rather than a data-entry one.

Increase accuracy β€” enhance existing items​

Finds additional reference images for items already in your catalogue that have too few references to be recognised reliably.

Use this for:

  • Items with low recognition accuracy and a thin reference set
  • Products that have been rebranded or repackaged, where existing references now look wrong
  • Broadening an item's references across store formats and lighting conditions

How it works​

  1. As your images are processed, the visual signatures (embeddings) of detected products are retained along with the candidate matches, date, and crop location.
  2. RDF clusters those signatures to find groups of the same product.
  3. Groups are matched against your existing catalogue β€” a close match becomes proposed references for that item, and a group matching nothing becomes a proposed new item.
  4. Each group is sent to the Real Data Flow channel in the Catalogue Gatekeeper as a submission, with its clustered crops attached.
  5. You review each submission and approve or reject it.

Only reasonably-sized, confident clusters are proposed β€” a group needs roughly 10 or more images before it is considered meaningful. This deliberately trades coverage for precision: a cluster built from three crops is far more likely to be noise than a real product.

info

RDF is triggered by Neurolabs rather than from the web app. Runs can be scheduled per organisation, or triggered on request against a specific task or date range. Talk to your Neurolabs contact about setting up a cadence β€” the right frequency depends on how many images you send.

Reviewing RDF submissions​

RDF proposals are machine-generated and need real scrutiny at review time. Two failure modes are worth knowing about:

Products that differ only by size. Two variants that are visually near-identical apart from a size printed on the pack will often land in the same cluster. Check the crops in a proposal actually belong to one variant before approving it as one item.

Mixed clusters. Occasionally a cluster contains more than one product. Reject rather than approve-and-fix β€” an approved mixed cluster teaches the models that two different products look the same.

warning

Approved reference images take effect immediately for subsequent images. Because RDF can generate a large number of submissions in one run, it is worth reviewing a sample carefully before working through the batch at speed, so you can spot a mis-tuned run early.

Clusters can also be tuned β€” how tightly grouped a cluster must be, and how aggressively near-duplicates are collapsed. If a run produces consistently loose or consistently fragmented clusters, that is a configuration fix rather than something to work around at review time; report it rather than rejecting hundreds of submissions.