Skip to contentv0.1.0
PIMgate

// Blog

10 practical tips for Open Data with industrial product data

How to publish industrial product data that machines can actually use: selection rules, units, identifiers, licensing, versioned feeds and change logs.

7 min readPIM Gate teamE-commerce

OPEN DATAGoverned catalogselection rulesOpen DataJSON-LD · FEEDS · API · LLMS.TXTSearch enginesMarketplacesAI agentsPUBLISH WHAT MACHINES CAN READ

Type one of your product names into an AI assistant and read the specification it gives back. If the measuring range is one you changed two years ago, you are not looking at a search problem — you are looking at your own data, copied by a reseller, outdated, and now treated as authoritative because it was the only machine-readable version available.

Open Data is the deliberate answer: publish a governed slice of your catalog as structured, public, machine-readable data so that the copy machines read is yours. Below are ten things that decide whether that publication is useful or merely present.

1. Decide what public means before you decide what to publish

Public is a governance question, not a technical one. Before any format discussion, get a written answer to: which products, which attributes, which languages, and who signs off.

Do: Draft a one-page Open Data profile. Product scope by category and lifecycle status; an explicit attribute allow-list; the languages; the owner. Everything not on the list stays entitled. Default to private — nothing becomes public because it happens to exist in the PIM.

2. Express the scope as rules, not as a list

A hand-picked list of 6,000 SKUs is correct on the day it is made and wrong by the next product launch. Rules survive the catalog.

Do: Write your selection as conditions over attributes you already maintain — category IN [Pressure, Temperature, Level], lifecycle = active AND market IN [DE, AT, CH], attribute.visibility = internal → exclude. Then check the resulting record count before publishing. If a rule change swings the count by thousands, you have found an attribute nobody maintains.

3. Use an attribute allow-list, never a block-list

A block-list fails the first time someone adds a new internal field. The list of things you forgot to exclude grows silently; the list of things you chose to include does not.

Do: Enumerate the public fields positively. In a typical industrial catalog: name, description, category, dimensions, technical ranges, materials, GTIN, images, data sheet. And positively exclude the categories of data that should never leave — list and customer prices, stock, cost, internal notes, supplier data, unreleased products.

4. Publish units as codes, not as strings

"10 bar" is a string a machine has to guess at. minValue: 0, maxValue: 10, unitCode: "BAR" is a fact it can compare, convert and filter.

Do: Map every numeric attribute to a value, a unit code (UN/CEFACT for schema.org contexts) and — where it applies — a min and max rather than a rendered range. Do this once in the canonical model and every output inherits it. If your PIM stores "0-10 bar" in a text field, that is your first data-quality ticket, and it will pay for itself again in ETIM and in the Digital Product Passport.

5. Give every product a stable, resolvable identifier

Machines reconcile records by identifier. If your public URI changes when a product moves category, every consumer that cached it now holds an orphan.

Do: Publish sku, mpn and gtin13 where you have them, and mint a stable @id URI that does not encode the category path. Keep it stable across renames and re-categorisations. Where GTINs exist, GS1 Digital Link syntax (/01/{GTIN}) is the form retail and scanning systems expect; it is planned for the module rather than available today, but designing your identifiers so it can be added later costs nothing now.

6. Use schema.org/Product as the vocabulary, and additionalProperty for the technical truth

schema.org is the established vocabulary search engines use to read product pages. Its core fields cover commercial facts; industrial specifications live in additionalProperty.

Do: Emit one JSON-LD document per product, embedded in your own product page (script tag or server-side include) and retrievable from an endpoint. Put every technical attribute in an additionalProperty array as a PropertyValue with name, value or minValue/maxValue, unitCode and unitText. Attach the data sheet as a subjectOf DigitalDocument with its encodingFormat and URL, so the PDF is discoverable as a document about the product rather than an unlabelled link.

7. Validate before you publish, and hold back what fails

Public data is quoted back at you. A record with a missing unit or a truncated description is worse in public than absent, because it will be mirrored.

Do: Define completeness rules for the Open Data profile specifically — stricter than your internal minimum. Records that fail are held, not published with gaps. Look at the held set weekly: it is the most honest data-quality backlog you will get, because every item on it is something a customer would have seen.

8. Version every release and keep a change log

"The feed is updated nightly" tells a consumer nothing they can act on. A version number and a change summary tell them whether to re-ingest.

Do: Number each publication run, publish added / changed / withdrawn counts with record IDs and timestamps, and keep earlier versions retrievable for a retention period you state. Ship a full feed and a delta feed side by side — marketplaces and wholesalers with large catalogs will take the delta and stop re-ingesting a million records to find 112 changes. Field-level delta sync is what makes that cheap: about 95 % smaller payloads than full exports.

9. Withdraw products explicitly — never delete them silently

A record that vanishes is indistinguishable from a fetch failure. Consumers cope with that by keeping the last copy they saw, which is exactly the outdated specification you are trying to eliminate.

Do: Mark discontinued products as withdrawn in the feed, with a date, and keep them resolvable at their URI for a defined period. If a successor exists, reference it. This is also the cheapest thing you can do for the AI-assistant case: an assistant that can see a product is withdrawn will say so.

10. State the licence, and make it machine-readable

Open Data without a licence is data nobody's legal department will let them use. Marketplaces and partners need to know what they may do with it and whom to credit.

Do: Set a licence per dataset — an open licence such as CC BY 4.0, or your own terms of use — and carry it in the feed manifest and the endpoint index, not only in a footer. Include the attribution string. The choice is yours to make with your legal team; the requirement is that it is explicit and retrievable.

An AI assistant quotes a measuring range we changed two years ago — from a reseller's copy of our data.

What this looks like operationally

Open Data should not be a second export pipeline. It is a published view of the same validated data your portals and APIs already serve, filtered by the rules in tip 2 — which means a specification change reaches your customer portal, your API consumers and the public feed from one correction, in one run.

A few operational points that come up every time:

  • Cache the public responses and invalidate on publication, so crawlers and partners never reach your source systems.
  • Read-only, with rate limits. A public endpoint has GET and nothing else.
  • Log rule changes with user and timestamp. "Who made this public" is a question that eventually gets asked.
  • Publish a product-data sitemap — a machine-readable index of public product URLs with last-modified dates — so crawlers do not have to discover your catalog by guessing.

The Open Data module is planned for 2026; GS1 Digital Link resolution and an llms.txt-style agent manifest are planned beyond that. None of the ten tips above depends on the module, though. Units as codes, stable identifiers, an attribute allow-list and a written profile are work in your own data model, and they are the part that takes the longest.

Key takeaways

  • Start with a written Open Data profile — scope, attribute allow-list, languages, owner — and default everything else to private.
  • Express scope as rules over maintained attributes, not as a hand-picked SKU list; check the record count before each publication.
  • Numbers need unit codes and min/max, not rendered strings; this single change pays out again in ETIM and DPP work.
  • Version every release, ship full plus delta feeds, publish a change log, and mark withdrawn products instead of deleting them.
  • Carry a licence and attribution in the feed manifest itself, so consumers can act on it without a legal enquiry.
  • Validate against a stricter profile for public data and hold back failures — the held set is your most honest quality backlog.

Put a gate between your catalog and chaos.