Skip to Content
SourcesLineage & Relationships

Lineage & Relationships

Assets rarely stand alone. A report is built from a dataset, which is built from a warehouse table, which was loaded from a file someone dropped in a bucket. A ticket has an attachment. A page mentions another page.

Lineage is the answer to two questions that keep coming up once you find something sensitive:

  • Where did this come from? — the customer export has card numbers in it; which upstream table leaked them in?
  • What breaks if I change it? — we want to drop this column; what stops working?

Classifyre records those relationships while it scans, from the systems’ own catalogs, and shows them on every asset.


Not every connection is lineage

The reason lineage graphs turn into unreadable hairballs is that everything gets flattened into one kind of “link”. An attachment, a foreign key, a duplicate file and a derived table are four completely different relationships, and only one of them answers “what breaks if I change this?”.

So every relationship Classifyre records carries a class that says what traversing it actually means:

ClassMeansExampleIs it lineage?
Lineage (FLOW)The values in one came from the otherA view over its base tablesYes — this is the lineage graph
Contains (CONTAINMENT)One is a part of the otherA chart in a dashboard; a file in an archiveNo — it’s what the graph is collapsed by
Same as (IDENTITY)The same real thing, seen twiceThe same file in two bucketsNo — it’s what nodes are merged by
References (REFERENCE)One points at the otherA foreign key; a page linking to a pageNo — it propagates nothing
Used by (USAGE)Somebody touched itA person who owns or opened itNo — it’s a weight, not a path

This distinction is what keeps the lineage view honest. A foreign key moves no data, so it is recorded — it’s a good hint about where lineage might exist — but it never becomes a hop in a lineage path. Everything the system is unsure about is filed as a reference, the class that propagates nothing, rather than being quietly admitted into lineage.


Kinds of lineage

Within lineage itself, the nature of the derivation is recorded too. In the app these read as plain phrases on the arrow:

Shown asWhat it means
derivedComputed from the upstream, with real logic in between
view ofA view or a report over its base tables
copy ofA replica or mirror — the same values, moved
written byA process or job produced this
exported toThe data left the system
sent toThe data was delivered to a recipient

Arrows always point the way the data moves: upstream on the left, downstream on the right. So walking outward from an asset answers “what breaks if this changes”, and walking inward answers “where did this come from”.


Column-level lineage

Table-to-table lineage tells you two tables are connected. That is often not precise enough to act on — before you drop a column you need to know whether this column feeds anything.

Where the information exists, Classifyre records the mapping per column: which upstream columns feed each output column, and the expression in between. Open an asset with columns and pick one to see exactly what feeds it.

A column can also matter without feeding any output column — a date used in a WHERE, a join key, a sort. Those shape which rows came out but no particular value, and they’re listed separately as indirect dependencies rather than being drawn as an arrow into a column they don’t actually feed.

Where column detail comes from. Some systems answer the column question directly (Databricks Unity Catalog, SQL Server). For the rest, it is recovered from the view’s own SQL. A SELECT * view yields no column mappings at all — the edge stays at table level, which is the honest answer. A confidently wrong column mapping is worse than a missing one.


How much to trust an edge

Not all lineage is equally certain, so each relationship records how it was derived. The app shows this next to the relationship when you select it:

How it was knownMeaning
Runtime observedSomething watched the data actually move
System catalogThe platform’s own catalog said so
SQL parsedRead out of the query or view definition
HeuristicInferred — a reasonable guess, no more
ManualA person drew it

Two sources are allowed to disagree about the same pair of assets; both edges are kept with their own provenance rather than one silently overwriting the other. Where an edge came from SQL, the SQL itself is kept with it, so you can read the derivation instead of trusting a label.


Lineage across systems

The most valuable lineage crosses a boundary — a Tableau data source reading a Snowflake table, a Databricks table loaded from an S3 path. The awkward part is that the two halves are usually scanned by different connectors, sometimes months apart.

Classifyre handles this by naming objects the way their own platform names them, independently of who scanned them. A Tableau scan can point at a Snowflake table it has never seen, and the edge is kept in place. If a Snowflake source is connected later, the two halves recognise each other and the lineage completes itself — retroactively.

Until the other system is connected, those endpoints appear in the lineage view as not yet scanned — outlined nodes with a name but no content. They’re a useful finding in themselves: they tell you which systems your data flows through that Classifyre isn’t watching yet.


Lineage as evidence in duplicate review

Lineage has a second job beyond “where did this come from”. It is the cross-check that makes duplicate review worth reading.

When two assets look nearly identical, the interesting question is not how similar are they but why. Duplicate matching answers that from the contents of the assets; lineage answers it from a completely different direction — connector catalogs, query logs, view SQL. Because the two do not share a source, combining them genuinely adds information.

Similar, and…Reads asWhat it meansPriority
there is a pathShared upstreamA derived copy — one comes from the other, or both from the same placeLow. Expected redundancy; a mart that resembles its source is doing its job.
there is no path, and both sides have lineageNo pathConvergent duplication — two teams built the same thing and nothing connects themHigh. Expensive and invisible; nobody has a reason to look.
we have no lineage for one sideLineage unknownA coverage gapNeither. Judge on the values.

Reporting derived copies as if they were problems is the main reason metadata-only duplicate detection gets ignored, which is why the first row is deprioritised and the second gets the alarm colour in the app.

Which classes count as derivation evidence

Not the same set the lineage view walks. For this test, FLOW, CONTAINMENT and IDENTITY all count as evidence that two assets are related by derivation — containment and identity both imply a real relationship between the things, not just a pointer. REFERENCE and USAGE never count: a foreign key moves no data, and somebody opening a file says nothing about where it came from.

The one exclusion that matters

Classifyre’s own similarity edges are never counted as lineage evidence.

The duplicate engine writes its results into the same relationship store, classed as REFERENCE (similar, likely duplicate) and IDENTITY (identical content). If the derivation test walked those, every pair in the review queue would find a “path” — to itself — every pair would report as explained, nothing would ever be escalated, and the whole check would look like it was working.

So relationships produced by the duplicate engine are excluded by name. A connector-declared same as stays in: different source, genuine evidence. The test is only useful because it is answered by something other than the thing it is checking.

When lineage stops being usable

The test approximates “is there a path between these two assets” with “are they in the same connected component”, which is far cheaper. That approximation fails in one direction: a single hairball component makes everything look derived, and the queue would quietly escalate nothing.

So when one component swallows too much of the lineage graph, Classifyre reports lineage unknown rather than claiming a path — and says so on screen. It is the honest answer to a question the shape of the graph has made unanswerable. The guard stays off on small estates, where lineage is sparse and every edge matters.


Where to see it in the app

Open any asset and choose the Lineage tab.

ControlWhat it does
Upstream / Downstream / BothWhich way to walk — where it came from, what depends on it, or the full picture
CollapseRolls each item up into whatever contains it (tables into schemas, charts into dashboards) — how four hundred tables become twelve schemas without losing an edge
Not yet scannedCounts endpoints in systems you haven’t connected

Below the graph, the column lineage panel traces a single column.

Relationships also appear throughout investigations and in duplicate review — in the case graph, relationship types are grouped by class, so lineage is visually distinct from containment, duplicates, and references. Selecting any relationship shows its class, how it was derived, whether it carries column detail, and the SQL behind it where there is any.

You can also draw a relationship by hand in a case graph when you know something the systems don’t. Manual edges are marked as such, so they’re never mistaken for something a platform reported.


Which sources produce lineage

Lineage is only as good as what the underlying system is willing to tell us. It is not a setting you switch on — a source produces it if its platform exposes it.

SourceWhat it produces
PostgreSQL, MySQL, SQL Server, Oracle, Snowflake, HiveViews and the tables they read from, with column detail parsed from the view SQL. Foreign keys as references.
DatabricksUnity Catalog lineage, including true column-level lineage from the catalog itself
TableauData sources and the warehouse tables behind them; workbooks and projects as containment
Power BIReports and dashboards over their datasets, and the databases those datasets pull from
SQLiteForeign keys as references
Custom connectorsWhatever you declare — every relationship class is available to you
Augmentation on any sourceWhatever your augment() yields — the same relationship classes, attached to a known source’s assets

Every other source still records links between the things it finds — a comment on its issue, an attachment on its page, a file to the message it was shared in. Those are relationships, and they show up in the graph; they are just not lineage, because no data moved.

Nothing to see yet? Lineage appears after a scan of a source that can report it. If an asset’s Lineage tab is empty, either its source doesn’t expose lineage, or the objects around it haven’t been scanned yet.


Building lineage yourself

If you know how your data moves and no system will tell us, a custom connector can declare it directly — every class on this page, column mappings included, plus the ability to point at objects in systems the connector doesn’t scan. See the notebook reference.

Last updated on