// Blog
Translating product data: delta, TM, terminology, and AI
Why catalog translation costs more than it should, and the mechanisms that fix it: segment hashing, deduplication, memory, termbase — with AI in its place.
8 min readPIM Gate teamMarketing & localization
You changed one attribute on eight hundred products and received a translation quote for forty thousand words. Somewhere between your PIM and the agency, a mechanism that should have recognised one changed sentence turned it into a file-sized job — and nobody in the chain can point to where.
The reason is almost always the same: the pipeline translates files when it should translate segments. Everything below follows from that one distinction.
The unit of work is a segment, not a file
A file-based workflow exports a product, a language or a catalog, hands it over and gets it back. It cannot know that of the 40,000 words in the export, 39,300 are identical to last month and 600 of the rest are one sentence repeated across a product family.
A segment-based workflow treats each attribute value and each sentence as an item with an identity. Give each one a hash computed over its normalized source text plus its context — attribute name, product type, locale — and three useful things become possible at once:
- Change detection. When a record comes through, only segments whose hash changed enter the pipeline. Editing one product never triggers re-translation of another.
- Deduplication. "Messing, vernickelt" appearing on 812 products is one segment. Translated once, reviewed once, applied 812 times.
- Traceability. For any segment you can answer where it is used, when its source last changed, and what its status is per target language.
The order of magnitude matters more than the exact figures. A catalog of 12,000 SKUs with 14 translatable attributes has around 168,000 field values; after deduplication, typically tens of thousands of unique segments; a weekly run changes maybe a thousand; after memory lookup a few hundred need genuinely new translation. That is the difference between a five-figure word count and an afternoon of review — and it comes from bookkeeping, not from AI.
One refinement worth configuring: whether a source edit touching only punctuation or whitespace counts as a change. Unchecked, a tidy-up pass by a product manager can invalidate a month of translations.
Translation memory: the asset your agencies already built for you
Most companies have years of TMX sitting with their agency and no way to use it against their catalog. Importing it changes the arithmetic immediately, because catalog language repeats heavily across products and across years.
The mechanics worth insisting on:
- Exact and fuzzy with thresholds you set. Exact matches are reused directly. Fuzzy matches — say 87 % — go to the engine or the translator as a starting point, not as a finished answer.
- Context-aware matching. The same German phrase needs different English in a short description than in a technical attribute. A memory that ignores context produces confident wrong answers, which are worse than none.
- Growth. Every approved translation writes back into the memory. The asset compounds; run 20 is cheaper than run 2.
- Portability. Insist on TMX export. A memory you cannot take with you is a lock-in, and it is your linguistic asset, not your vendor's.
Terminology: enforced, not suggested
A glossary in a shared document is advice. A termbase in the pipeline is a control.
Import your approved terms as TBX, with three states per term and language: approved, preferred, forbidden. Then two things happen at different points in the pipeline, and both are necessary:
- Before translation. Approved terms are passed to the engine as a glossary where the engine supports it.
- After translation. Every proposal is checked against the termbase regardless of which engine produced it — because a glossary is a hint to a model, not a guarantee.
A violation then either blocks approval or raises a warning, per language, by your choice. "Manometer" where your termbase says "pressure gauge" never reaches a reviewer unflagged, and never reaches a channel at all.
Two things repay the setup effort: locale-specific entries that override the base language, and a route for reviewers to propose terms from the queue, so the termbase grows from real work rather than from a one-off workshop.
Where AI actually fits
AI translates. It does not decide. Keeping those separate is what makes it safe to use on a catalog.
In practice, machine translation produces a proposal that enters the review policy you configured for that language. Different languages deserve different policies: a high-volume language with a mature memory might be machine proposal plus single review; a legally sensitive market, human translation only, with the pipeline producing packages rather than proposals; a locale variant like de-CH or en-US is better derived from its base language by rules — ss for ß, Swiss terms, number format — than translated from scratch.
Engine choice belongs per language, not per company. Quality varies by language pair and content type, so connect DeepL, OpenAI, Anthropic, Azure OpenAI or an EU-hosted or local model, and change one without touching the workflow around it. That also keeps the data-location question answerable: if a market requires it, route that language to a model hosted where you need it.
A reviewer opens a segment and sees the source, the proposal, the 87 % memory match it came from, the term hits inside it, and the channel length limit it has to respect — on one screen, before deciding.
That screen is the actual product of the pipeline. Everything upstream exists to make the reviewer's decision small and well-informed.
The constraints nobody remembers until a channel rejects the file
Translation is not finished when it reads well. It is finished when every channel accepts it. Four checks belong in the pipeline rather than in the recipient's error log:
- Length limits per attribute and channel — a marketplace title capped at 80 characters, a label field, a table cell in a PDF. French and German expansion breaks these routinely; flag over-length before review, not after publication.
- Placeholders and markup — variables like
{size}, line breaks, inline tags — protected during translation and verified on return. - Units, standards and article numbers — "G 1/4", "DN 15", "EN 837-1" — locked so no engine helpfully localizes them.
- Forbidden content per channel, caught at the gate rather than by the channel.
Keep your translators where they are
The most common objection to any translation tooling is that the agency will not work in it. They do not have to. Export a package — only new and fuzzy segments for one language, deduplicated, with context, memory matches and terminology attached, as XLIFF or XLSX — and the agency works in their own CAT tool with no account in your system. The package comes back, is validated on import (segment IDs, placeholders, lengths, terminology, and whether the source changed since export), and enters the same review queue as any other proposal.
Say the commercial effect plainly to your agency: they quote on real work, not on repetitions. That conversation goes better when you can show a package with the duplicate count at zero.
What to do first
Start with an inventory, not a tool decision. Count translatable attributes and multiply by SKUs. Ask your agency for the TMX files. Write down which terms have ever caused an argument — that is your first termbase. Then take one target language end to end.
See the pipeline in detail, including memory, termbase, engine routing and review queues. Product-data translation ships as a module on the same data quality gate as everything else; content translation connectors for Ibexa DXP and Drupal are planned and currently in early access, so plan around the product-data workflow first.
Key takeaways
- Translate segments with hashes, not files: change detection, deduplication and traceability all follow from that one change.
- Import the TMX your agencies already built; require context-aware matching and TMX export so the memory stays yours.
- A termbase must be enforced twice — passed to the engine, then checked on the result, regardless of engine.
- AI produces proposals; the review policy per language decides what publishes. Derive locale variants by rules rather than re-translating.
- Choose the engine per language, and keep it swappable so data-location requirements stay answerable.
- Check length, placeholders, locked units and forbidden content in the pipeline, not in the channel's error log.
- Export deduplicated delta packages so agencies stay in their CAT tool and quote on real work.
// Related articles
8 min readIT & architecture
Your portals should outlive your PIM contract
PIM-native portals die with the PIM. A canonical layer plus a migration bridge lets you change the system underneath without touching a single channel.
architecturemigrationitportals
6 min readProduct data
Why you need a customer portal when your PIM already exports
Exports are a delivery mechanism, not an access model. What a customer portal adds: entitlements, live data, self-service and a record of every download.
portalsentitlementsproduct-data