The phase that consumes an ETL project, run by an agent
Data unification projects stall on one question: what does this data mean? A column named CX138 in a legacy table has no meaning on its own, and the documentation or the person who could explain it is often hard to locate. The gathering agent assembles that context across your systems and asks the people who hold the rest of it.
Assembling meaning is the majority of the work
Reading an ERP system, a time-series historian and a handful of departmental applications well enough to unify them is not a technical problem so much as an archaeological one. It typically consumes the majority of an ETL engagement, and it is the part no connector catalogue shortens.
Locate, scan, flag, ask
A gathering manager on your data team initiates discovery for a set of target sources. The agent works through four steps and reports back on each.
Locate
The agent identifies where the relevant datasets reside and who administers them, then guides those administrators through connecting each source.
Scan
SetMeld Composer scans each connected source for tables, columns, types, keys, sample values and relationships, establishing the system’s understanding of it.
Flag
The scan writes a description of every element. Whatever it cannot settle from the data alone is flagged as needing a person, rather than guessed at.
Ask
The agent finds the subject matter expert for each flagged element and sends them a targeted question over email or Slack. Every answer is recorded against the element it clarifies.
Any structured source, no prebuilt connector
Relational databases, document stores, REST APIs and time-series historians are all in scope. The scan is what establishes the system’s understanding of a source, which is precisely why a prebuilt connector for it is unnecessary: nothing is being pattern-matched against a vendor template.
- Structure and semantics together. Tables, columns, types, keys, sample values and relationships, each with a description of what it appears to mean.
- Value distributions are evidence. A column whose values contradict its apparent purpose is caught here rather than three weeks into the mapping.
- Documentation is an input. Data dictionaries, schema notes and internal wikis uploaded alongside a source reduce how much has to be asked.
- Most columns never need a person. On the project pictured below, 24 of 86 scanned columns did. The scan settled the rest on its own.
A question to one person, about one column
Ambiguous column names, acronyms with several plausible meanings, and value distributions that conflict with a column’s apparent purpose all become questions addressed to a named owner. What follows is one of them, end to end.
The question leaves the building
The agent works out who to ask and writes to them where they already read their messages. The mail names the exact column, explains in plain language what is unknown about it and what cannot be computed until it is settled, and links straight to that one question.
- Addressed to a person, not a distribution list. The expert is identified from who administers the system and who has been named as knowing the field.
- Context, not a ticket. The message says what SetMeld already inferred and why the answer matters, so the expert can judge it without opening anything.
- No tool to learn. Email or Slack, whichever they use. There is no account to request and no product to be trained on.
Confirm it, correct it, or hand it on
The link opens one question against one column, with the meaning SetMeld proposed and the schema it sits in. Agreeing, correcting it, or passing it to somebody who would know better all happen there, and an answer written as a sentence is as good as a button.
- Three ways out, all of them cheap. That’s right, Correct it, or Ask someone else. Nobody is stuck on a question they cannot answer.
- The column in context. It is highlighted in the live schema, with its table, its type and its neighbours, so the expert can see what is being asked about.
- Recorded, not merely replied to. The advisor writes the settled meaning back against the column and says so, so the expert can see their answer landed.
You can see how much context you actually have
The gathering manager watches one board: the percentage of context gathered, the questions outstanding, and every open column beside what SetMeld thinks it means and who it is with. The phase is complete when collected context reaches the confidence threshold required for a high-quality knowledge graph, not when someone declares the workshops finished.
- Three states, not a progress bar. Confirmed, with somebody, or no one to ask. Only the last of those needs a manager.
- Sources and people both have owners. Each datasource shows what is settled and who it waits on; each person shows how many questions sit with them, and for how long.
- A stalled source is obvious. Coverage that stops moving points at a specific unanswered question and a specific name, rather than at a status meeting.
“ABS” means two different things in two tables
Enterprise data is dense with context-dependent abbreviations. ABS may be the American Bureau of Shipping in one table and a plastic in another, and no amount of schema inspection settles which. The agent resolves these from two kinds of source.
The people who know
Subject matter experts answer in their own words, attach the documentation they have, or point the agent at the colleague who actually owns the answer.
Reference material
Internal glossaries, data dictionaries and schema notes, with the language model proposing candidate meanings for the expert to confirm or reject.
Ambiguity stays visible
Anything still unresolved after both sits under no one to ask, rather than being quietly settled by a guess that surfaces later as a wrong join.
What the gathered context becomes
The output is a structured knowledge base describing every table, column, key and relationship, annotated with expert clarifications. SetMeld Composer feeds it into the AI pipeline that designs the integration.
Map entities
An alarm event in a historian is linked to its tag, and the tag to the equipment record in the ERP.
Generate the ontology
The unifying ontology and target schema are derived from what was gathered, not from a template.
Produce the ETL configuration
Extraction plans, transformation logic, entity resolution rules and load specifications, all reviewable.
Run it
SetMeld Pipeline executes that configuration, cleansing, deduplicating and resolving entities, and loads the unified knowledge graph.
Gathering questions
Does anything leave our environment?
Do you need a connector built for our historian first?
How much of our experts’ time does this take?
What if nobody answers, or nobody knows?
Who runs the process on our side?
Point the gathering agent at your least-documented system
The one with the column names nobody can explain is the useful test, not the tidy one.