Faceted search
Filters that actually narrow things down, because there is a maintained field behind every one of them.
Product lists, course catalogues, job ads, provider directories: usually everything is there, sitting in free text, written differently by every supplier and missing exactly the attributes a customer would filter by. The pipeline normalises, classifies, enriches and tags it, record by record and traceably. The work runs on a local language model, on your server or on ours, and no record is sent to a cloud provider.
Business English, 20 lessons/week à 45 min, small group max. 8, from B1, Malta, 2 weeks from 690.- incl. material
The price per week is the interesting one. Nobody wrote it down, and no visitor can compare offers without it.
Every stage does one job and can be checked on its own. Running them together is what makes results impossible to debug later.
Units, spellings, currencies and date formats are brought into one shape.
From 690.- and EUR 690.00 comes the same number. From 20 lessons and 20 teaching units comes the same figure.
Every record gets its category, sub-category, audience and level.
The model does not invent categories, it picks from your list. A record that fits nowhere goes on the review list instead of into the nearest box.
What sits in the text but not in a field is pulled out and stored separately.
Group size, minimum age, what is included, what is excluded. These are the attributes people filter by and almost nobody maintains.
Facets for the filter navigation and vectors for similarity search.
Afterwards somebody can find a small group course for beginners by the sea without any of those words appearing in the record.
Refining data means sending every record through a model, several times, and doing it again whenever the field list changes. That is exactly the workload where the cloud gets expensive and uncomfortable.
| Cloud API | Local model | |
|---|---|---|
| Where the data goes | Every record is transferred to a provider, often outside the EU | The data stays on the machine it already lives on |
| Cost at volume | Per record, every time, including the second run | Compute time, with no price per record |
| Repeatability | A model update at the provider can change results | Model and version stay fixed until you change them |
| Traceability | A value arrives with no way to see where it came from | Every value carries the text passage it was taken from |
| Speed on large stocks | Capped by the provider's rate limits | As many parallel runs as the hardware carries |
| Confidentiality | Prices, margins and suppliers leave the building | Nothing leaves the building |
The same hardware runs your other AI workloads KI-Server
Refining is never the goal. It is the step that has to happen before any of the following actually works.
Filters that actually narrow things down, because there is a maintained field behind every one of them.
Facets become indexable pages of their own. Twenty-six locations and eight categories produce pages that did not exist before.
A generated brochure can only contain matching offers if the matching is in the data. Brochure on Demand
Price per week, per unit, per participant. Only normalised figures can be compared at all.
Semantic hits instead of exact word matches, using vectors calculated locally.
What is missing shows as not stated, never as zero. That difference decides every evaluation built on top.
A language model asked to fill a gap will fill it, if necessary with something plausible. On product data that is the most expensive mistake available, because an invented price does not look like a mistake. So the pipeline has one hard rule: every generated value carries the passage it came from. If there is no passage, the field stays empty and the record goes on a review list.
A named gap beats a guessed number. It costs an afternoon of manual work and saves the phone call from the customer who booked at a price that was never real.
Business English, 20 lessons/week à 45 min, small group max. 8, from B1, Malta, 2 weeks from 690.- incl. material
Review list for uncertain values
The packages differ in how many records you refine. Everything else is in all of them.
FitsOne product range or one directory.
Send a sample fileFitsA portal with continuous new entries.
Send a sample fileFitsMarketplaces and aggregators with feeds from many suppliers.
Send a sample fileRuns on our server or on an AI server of your own. Above roughly fifty thousand records your own hardware is usually cheaper.
An open model on our own hardware, chosen for your case, plus a separate embedding model for similarity search. Both are named in the documentation so a run from today can still be reproduced in a year.
On tidy source text the hit rate per field is high. On free text written by 40 suppliers over 15 years it is not, which is what the review list is for. We measure against a sample you fill in by hand and tell you the rate before you decide.
They run through the same pipeline automatically. A new import does not require a second pass over the old stock.
Yes, and the correction wins. A value set by hand is never overwritten by a later run.
For small stocks no, it runs on ours. From a few tens of thousands of records your own hardware pays off, and that is where our AI server fits in.
German and English are proven. Other languages are a question of the model, not of the pipeline.
Ten thousand records are usually through overnight, depending on how many fields you want. Setting the field model up beforehand takes longer than the run itself.
You keep the refined data, the field model and the scripts. All of it sits in your database, not in ours.
A slice of your stock, however untidy. You get the same fifty rows back refined, together with the review list and the honest number of fields we could fill with confidence.
No obligation, and the file never leaves our machine.