Gathering agent

The phase that consumes an ETL project, run by an agent

Data unification projects stall on one question: what does this data mean? A column named CX138 in a legacy table has no meaning on its own, and the documentation or the person who could explain it is often hard to locate. The gathering agent assembles that context across your systems and asks the people who hold the rest of it.

The bottleneck

Assembling meaning is the majority of the work

Reading an ERP system, a time-series historian and a handful of departmental applications well enough to unify them is not a technical problem so much as an archaeological one. It typically consumes the majority of an ETL engagement, and it is the part no connector catalogue shortens.

SetMeld automates this phase with a gathering agent. It runs inside SetMeld Composer, deployed behind your firewall, and can use your own language model endpoint, so no data leaves the environment. See the deployment models.
The workflow

Locate, scan, flag, ask

A gathering manager on your data team initiates discovery for a set of target sources. The agent works through four steps and reports back on each.

1

Locate

The agent identifies where the relevant datasets reside and who administers them, then guides those administrators through connecting each source.

2

Scan

SetMeld Composer scans each connected source for tables, columns, types, keys, sample values and relationships, establishing the system’s understanding of it.

3

Flag

The scan writes a description of every element. Whatever it cannot settle from the data alone is flagged as needing a person, rather than guessed at.

4

Ask

The agent finds the subject matter expert for each flagged element and sends them a targeted question over email or Slack. Every answer is recorded against the element it clarifies.

Scan

Any structured source, no prebuilt connector

Relational databases, document stores, REST APIs and time-series historians are all in scope. The scan is what establishes the system’s understanding of a source, which is precisely why a prebuilt connector for it is unnecessary: nothing is being pattern-matched against a vendor template.

  • Structure and semantics together. Tables, columns, types, keys, sample values and relationships, each with a description of what it appears to mean.
  • Value distributions are evidence. A column whose values contradict its apparent purpose is caught here rather than three weeks into the mapping.
  • Documentation is an input. Data dictionaries, schema notes and internal wikis uploaded alongside a source reduce how much has to be asked.
  • Most columns never need a person. On the project pictured below, 24 of 86 scanned columns did. The scan settled the rest on its own.
1 · Connect Data Sources
SetMeld Composer connecting sources and reporting scan results in real time
Flag, ask, settle

A question to one person, about one column

Ambiguous column names, acronyms with several plausible meanings, and value distributions that conflict with a column’s apparent purpose all become questions addressed to a named owner. What follows is one of them, end to end.

Ask

The question leaves the building

The agent works out who to ask and writes to them where they already read their messages. The mail names the exact column, explains in plain language what is unknown about it and what cannot be computed until it is settled, and links straight to that one question.

  • Addressed to a person, not a distribution list. The expert is identified from who administers the system and who has been named as knowing the field.
  • Context, not a ticket. The message says what SetMeld already inferred and why the answer matters, so the expert can judge it without opening anything.
  • No tool to learn. Email or Slack, whichever they use. There is no account to request and no product to be trained on.
The expert’s inbox
The expert’s inbox: an email from SetMeld naming the ad_exposure.exposure_code column, explaining what is unknown about it and linking to the question
Settle

Confirm it, correct it, or hand it on

The link opens one question against one column, with the meaning SetMeld proposed and the schema it sits in. Agreeing, correcting it, or passing it to somebody who would know better all happen there, and an answer written as a sentence is as good as a button.

  • Three ways out, all of them cheap. That’s right, Correct it, or Ask someone else. Nobody is stuck on a question they cannot answer.
  • The column in context. It is highlighted in the live schema, with its table, its type and its neighbours, so the expert can see what is being asked about.
  • Recorded, not merely replied to. The advisor writes the settled meaning back against the column and says so, so the expert can see their answer landed.
Gathering · Datasource SME
The expert’s view of a single question: the proposed meaning of exposure_code, the column highlighted in the schema, and the options to confirm, correct or reassign it
Completeness

You can see how much context you actually have

The gathering manager watches one board: the percentage of context gathered, the questions outstanding, and every open column beside what SetMeld thinks it means and who it is with. The phase is complete when collected context reaches the confidence threshold required for a high-quality knowledge graph, not when someone declares the workshops finished.

  • Three states, not a progress bar. Confirmed, with somebody, or no one to ask. Only the last of those needs a manager.
  • Sources and people both have owners. Each datasource shows what is settled and who it waits on; each person shows how many questions sit with them, and for how long.
  • A stalled source is obvious. Coverage that stops moving points at a specific unanswered question and a specific name, rather than at a status meeting.
Gathering
The gathering board: 41 percent of context gathered, 14 questions outstanding, a breakdown of confirmed questions against those with somebody and those with no one to ask, and a table of every open column
Tribal knowledge

“ABS” means two different things in two tables

Enterprise data is dense with context-dependent abbreviations. ABS may be the American Bureau of Shipping in one table and a plastic in another, and no amount of schema inspection settles which. The agent resolves these from two kinds of source.

The people who know

Subject matter experts answer in their own words, attach the documentation they have, or point the agent at the colleague who actually owns the answer.

Reference material

Internal glossaries, data dictionaries and schema notes, with the language model proposing candidate meanings for the expert to confirm or reject.

Ambiguity stays visible

Anything still unresolved after both sits under no one to ask, rather than being quietly settled by a guess that surfaces later as a wrong join.

Downstream

What the gathered context becomes

The output is a structured knowledge base describing every table, column, key and relationship, annotated with expert clarifications. SetMeld Composer feeds it into the AI pipeline that designs the integration.

1

Map entities

An alarm event in a historian is linked to its tag, and the tag to the equipment record in the ERP.

2

Generate the ontology

The unifying ontology and target schema are derived from what was gathered, not from a template.

3

Produce the ETL configuration

Extraction plans, transformation logic, entity resolution rules and load specifications, all reviewable.

4

Run it

SetMeld Pipeline executes that configuration, cleansing, deduplicating and resolving entities, and loads the unified knowledge graph.

Sources can be added incrementally. Each new one runs through the same gathering process, and SetMeld Composer extends the existing ontology to accommodate it, rather than starting the model again. What the graph gives you.
FAQ

Gathering questions

Does anything leave our environment?
In a self-hosted deployment, no. The gathering agent runs inside SetMeld Composer behind your firewall and can be pointed at your own language model endpoint, so schema, sample values and expert answers stay in your network. See deployment and security for the SetMeld Cloud model, where scanning does send schema information and data samples to SetMeld’s environment.
Do you need a connector built for our historian first?
No. The scan is what establishes the system’s understanding of a source, so any structured source is supported: relational databases, document stores, REST APIs and time-series historians included. Prebuilt connectors are unnecessary because nothing is being matched against a vendor template.
How much of our experts’ time does this take?
Each question is about one column, is addressed to the one person likely to know, and is answered by confirming the proposed meaning, correcting it, or naming somebody better placed. There is no workshop series and nothing for an expert to learn: replying to the message is the whole interaction.
What if nobody answers, or nobody knows?
The question sits under no one to ask on the gathering board, with the source it is holding back and the days it has been open. Unresolved ambiguity is reported rather than resolved by assumption, which is the point: a guess here becomes a wrong join in the graph months later.
Who runs the process on our side?
A gathering manager on your data team initiates discovery for a set of target sources and watches the board. Administrators are pulled in only to connect their own systems, and subject matter experts only to settle questions about their own data.

Point the gathering agent at your least-documented system

The one with the column names nobody can explain is the useful test, not the tidy one.