Product · Data Refinery

Raw records in, searchable data out.

Product lists, course catalogues, job ads, provider directories: usually everything is there, sitting in free text, written differently by every supplier and missing exactly the attributes a customer would filter by. The pipeline normalises, classifies, enriches and tags it, record by record and traceably. The work runs on a local language model, on your server or on ours, and no record is sent to a cloud provider.

  • Local model, no cloud
  • Servers in Germany
  • Every value with its source
  • Repeatable runs
One record, before and after
How it arrives

Business English, 20 lessons/week à 45 min, small group max. 8, from B1, Malta, 2 weeks from 690.- incl. material

How it looks afterwards
Category
Language course, professional
Location
Malta
Intensity
20 lessons per week, 45 minutes each
Group size
maximum 8
Minimum level
B1
Price
€690 for 2 weeks, €345 per week
Included
Course material
Facets
Professional · Small group · Intermediate

The price per week is the interesting one. Nobody wrote it down, and no visitor can compare offers without it.

Four stages, in this order

Every stage does one job and can be checked on its own. Running them together is what makes results impossible to debug later.

  1. 1

    Normalise

    Units, spellings, currencies and date formats are brought into one shape.

    From 690.- and EUR 690.00 comes the same number. From 20 lessons and 20 teaching units comes the same figure.

  2. 2

    Classify

    Every record gets its category, sub-category, audience and level.

    The model does not invent categories, it picks from your list. A record that fits nowhere goes on the review list instead of into the nearest box.

  3. 3

    Enrich

    What sits in the text but not in a field is pulled out and stored separately.

    Group size, minimum age, what is included, what is excluded. These are the attributes people filter by and almost nobody maintains.

  4. 4

    Tag

    Facets for the filter navigation and vectors for similarity search.

    Afterwards somebody can find a small group course for beginners by the sea without any of those words appearing in the record.

Why this runs on your own model

Refining data means sending every record through a model, several times, and doing it again whenever the field list changes. That is exactly the workload where the cloud gets expensive and uncomfortable.

Cloud APILocal model
Where the data goesEvery record is transferred to a provider, often outside the EUThe data stays on the machine it already lives on
Cost at volumePer record, every time, including the second runCompute time, with no price per record
RepeatabilityA model update at the provider can change resultsModel and version stay fixed until you change them
TraceabilityA value arrives with no way to see where it came fromEvery value carries the text passage it was taken from
Speed on large stocksCapped by the provider's rate limitsAs many parallel runs as the hardware carries
ConfidentialityPrices, margins and suppliers leave the buildingNothing leaves the building

The same hardware runs your other AI workloads KI-Server

What you can build once the data is clean

Refining is never the goal. It is the step that has to happen before any of the following actually works.

Faceted search

Filters that actually narrow things down, because there is a maintained field behind every one of them.

Landing pages per combination

Facets become indexable pages of their own. Twenty-six locations and eight categories produce pages that did not exist before.

Brochures and catalogues

A generated brochure can only contain matching offers if the matching is in the data. Brochure on Demand

Comparison and ranking

Price per week, per unit, per participant. Only normalised figures can be compared at all.

Similarity search

Semantic hits instead of exact word matches, using vectors calculated locally.

Visible gaps

What is missing shows as not stated, never as zero. That difference decides every evaluation built on top.

What the model is not allowed to do

A language model asked to fill a gap will fill it, if necessary with something plausible. On product data that is the most expensive mistake available, because an invented price does not look like a mistake. So the pipeline has one hard rule: every generated value carries the passage it came from. If there is no passage, the field stays empty and the record goes on a review list.

A named gap beats a guessed number. It costs an afternoon of manual work and saves the phone call from the customer who booked at a price that was never real.

Price
€690 for 2 weeks, €345 per week

Business English, 20 lessons/week à 45 min, small group max. 8, from B1, Malta, 2 weeks from 690.- incl. material

Visible gaps

Review list for uncertain values

One price, three sizes

The packages differ in how many records you refine. Everything else is in all of them.

Setup
from €2,900 once
Operation
from €190 a month

Medium

Most chosen
Recordsup to 50,000

FitsA portal with continuous new entries.

Send a sample file

Large

Recordsunlimited

FitsMarketplaces and aggregators with feeds from many suppliers.

Send a sample file

In every package

  • Field and facet model built from your own stock
  • Review list for uncertain values
  • First run across the entire stock
  • New entries refined automatically
  • Export into your database, search or website
  • Hosting, updates and support

Runs on our server or on an AI server of your own. Above roughly fifty thousand records your own hardware is usually cheaper.

Questions before you send the first file

Which model is doing the work?+

An open model on our own hardware, chosen for your case, plus a separate embedding model for similarity search. Both are named in the documentation so a run from today can still be reproduced in a year.

How accurate is it?+

On tidy source text the hit rate per field is high. On free text written by 40 suppliers over 15 years it is not, which is what the review list is for. We measure against a sample you fill in by hand and tell you the rate before you decide.

What happens to new records?+

They run through the same pipeline automatically. A new import does not require a second pass over the old stock.

Can we correct the results?+

Yes, and the correction wins. A value set by hand is never overwritten by a later run.

Do we need a server of our own?+

For small stocks no, it runs on ours. From a few tens of thousands of records your own hardware pays off, and that is where our AI server fits in.

Which languages?+

German and English are proven. Other languages are a question of the model, not of the pipeline.

How long does the first run take?+

Ten thousand records are usually through overnight, depending on how many fields you want. Setting the field model up beforehand takes longer than the run itself.

What happens if we stop?+

You keep the refined data, the field model and the scripts. All of it sits in your database, not in ours.

Send us fifty rows.

A slice of your stock, however untidy. You get the same fifty rows back refined, together with the review list and the honest number of fields we could fill with confidence.

Send a sample file

No obligation, and the file never leaves our machine.