Your mainframe is not going to emit OpenLineage events

Bringing Attribute Lineage to IBM z/OS Mainframes

What runtime capture reaches, and what you have to model yourself.

OpenLineage is an open standard for collecting data lineage as events. As a job runs, an integration reports which datasets it read, what processing it performed, and which datasets it wrote. It graduated from the LF AI & Data Foundation in 2023, and integrations now ship for Apache Airflow, Apache Spark, Apache Flink, dbt, and a range of data warehouses.1

Support for the standard has spread far enough across governance platforms that it no longer separates them, which makes “do you support OpenLineage” a weak question to build an evaluation around. The more useful question is what the standard was designed to do, because that determines what it will never do for you, regardless of how many connectors ship.

Enterprise metadata has produced few standards that anyone adopts, and OpenLineage is a genuine exception. Understanding why it succeeded also explains what it leaves for you to solve.

What the specification defines

OpenLineage’s object model rests on three entities.2

  1. Job: “a process that consumes or produces Datasets.”
  2. Run: “an instance of a Job that represents one of its occurrences in time,” identified by a unique identifier.
  3. Dataset: an abstract representation of data with a name in a namespace.

As a job executes, it emits run events describing what it read and what it wrote.

Richer metadata attaches through facets, extensible JSON objects covering technical and operational detail such as schema, SQL, and column-level lineage. Organizations can define custom facets for anything the spec does not cover, and the documentation notes that “custom facets can be promoted to the standard by including them in the spec.”3 That extension path lets the standard grow without a central authority deciding what belongs.

Is OpenLineage runtime-only? The claim circulates widely, and it is wrong. Job events and dataset events carry design-time metadata, including declared inputs and outputs, schema, and source-code location.4 What limits most implementations is the producers, not the specification, since integrations like the Apache Spark listener emit when a job runs. The more accurate description of OpenLineage is an event-emission standard for technical lineage.

“Design lineage events (DatasetEvent, JobEvent) are not associated with a Run and represent design-time metadata.”

— OpenLineage, Object Model specification

A COBOL batch process moving balances overnight emits nothing, because there is no listener inside it and no supported framework reporting what it did. Every producer in wide deployment reads a modern engine, whether that is Spark, Airflow, dbt, Flink, or a cloud warehouse, and none of them reaches the mainframe. The standard does leave a second path open, since a tool that parses source code can emit design-time events describing what the code would do. That path asks you to analyze and model the code rather than instrument the runtime, which is a different discipline with a different cost, and almost nothing in the ecosystem does it today.

What the standard is genuinely good at

Inside a tool that emits, OpenLineage does something that no amount of documentation discipline achieves. It reports column-level lineage as a byproduct of the job running, without anyone maintaining a diagram. A Spark job that gains a new join produces events describing the new join the next time it executes, because the integration reads the execution plan rather than a description somebody wrote down.

That property solves the problem that defeats most manual lineage efforts. Handwritten documentation describes the pipeline as it was understood on the day someone wrote it, and pipelines change faster than the documentation about them. Lineage emitted by the engine cannot drift from what the engine did, since it comes from the same execution.

The format matters as much as the capture. Because the events are a published open specification rather than a proprietary export, lineage collected from Spark can be read by any tool that implements the standard, and the organization doesn’t have to negotiate an adapter for every pair of systems it owns. For teams standardizing on a modern pipeline stack, that portability is worth having on its own terms.

Both properties hold only within the same limit. They apply to the tools instrumented to emit, and to the flows that have already run. What follows comes from that scope rather than from any shortcoming in how the standard was built.

The omissions were deliberate, and they were the right call

The specification defines no business glossary and no linkage to governance policy, and it says nothing about version control or point-in-time history. Reading that list as a set of shortcomings misreads how standards get built.

When OpenLineage graduated, its Technical Steering Committee members were affiliated with Apache Airflow, Apache Iceberg, Apache Parquet, dbt, Egeria, Marquez, Microsoft, Snowflake, and Superconductive.5 Those nine organizations could agree on a schema describing what a job read and what it wrote, because that question has a mechanical answer any engine can report the same way. They could never have agreed on what a customer record means or who owns it. They would have disagreed again on which retention policy governs it, and on how to show that it meant something different in March. Those questions have institutional answers rather than mechanical ones, and every one of the nine would have answered differently.

A specification that tried to mandate business semantics would have died in committee, and those projects would have kept their proprietary formats. Narrow technical scope is why OpenLineage reached graduation, and why a market that agrees on almost nothing was able to adopt it. The tradeoff is straightforward: the scope that made the standard adoptable is the same scope that leaves everything else to whoever runs the data environment.

What the event stream looks like after a year

A pipeline running hourly emits roughly 720 events a month, every one describing the same unchanged data flow. Run that across a few hundred pipelines for a year and the volume becomes substantial, while the number of distinct flows it describes barely moves.

None of that indicates a defect. OpenLineage records executions because executions are what it was built to record. An architect looking for a map of the enterprise data landscape is asking a different question of the same material, and converting a record of ten thousand job runs into a model of how data moves is work that somebody has to do. Platforms that consume the standard differ mainly in how much of that conversion they perform for you, and how repeated runs get consolidated into one versioned flow is a fair thing to ask any of them to demonstrate.

Granularity compounds the conversion problem. The specification permits flexibility in what a producer includes, so some integrations provide column-level detail while others provide only table-level detail. Coverage across your pipelines ends up uneven in ways that depend on which framework emitted which event, and whatever consumes the stream has to reconcile that unevenness.

The sharper constraint shows up before deployment. A flow that has not run yet has emitted nothing, so a proposed schema change has no event history to analyze against. Teams that want to know what a change will break have to model the intended flow rather than wait for it to execute and read the wreckage afterward, which is the difference between impact analysis and incident review.

What a regulated firm still has to build

Supervisors under BCBS 239 and the Digital Operational Resilience Act (DORA) ask questions that a well-instrumented event stream cannot answer, and higher coverage percentages don’t resolve them.

Which systems feed this regulatory report, including the ones that never emit an event? Mainframe extracts, vendor feeds arriving as files, and the manual reconciliation a team performs at quarter end all move data that matters, and none of them appear in any event stream. Automated capture sets a ceiling, and modeling the rest is what gets you past it.

What does this data element mean, and who is accountable for it? Ownership, business definitions, and the policies attached to a field live in the institution rather than in the pipeline, and no facet emitted by Spark supplies them.

The hardest question concerns time. What did the lineage look like in March, before the schema change, when the report was filed? Answering requires the data lineage model as it stood on that date, carrying the governance context that applied then. A runtime log records which jobs ran in March, which tells you considerably less than how the data was governed then. What regulators expect is the state of the model, not the execution history. Version control over the lineage model and point-in-time reconstruction are what make that answerable, and neither belongs in an event-emission standard.

Where to start

Map what can emit before you evaluate anything. List which systems in your data environment produce OpenLineage events at all, and treat that list as the permanent ceiling on what automated capture will ever cover for you. Everything outside it needs a modeled representation, and buying more connectors will not shrink the list.

Test a past date rather than a current diagram. Pick a regulatory report, pick a date before your last significant schema change, and try to reconstruct what fed it and who owned each element. Whether you can answer tells you whether you hold a governed data lineage model or a well-organized log.

Stitch across one system that does not emit. Take a single Spark flow, connect it upstream to the non-emitting source that feeds it, and connect it downstream to the report it reaches. That exercise shows how much of your actual data ecosystem sits outside the instrumented middle, and it costs an afternoon rather than a procurement cycle.

Solidatus ingests OpenLineage events and turns them into the governed model those three tests ask for, harmonizing repeated runs into a versioned data lineage model, stitching that model to systems the standard never instruments, and keeping the point-in-time history an examiner asks for. The OpenLineage integration page covers how the connector works.

[1]LF AI & Data Foundation. “LF AI & Data Foundation Announces Graduation of OpenLineage Project.” LF AI & Data Foundation, September 20, 2023.
https://lfaidata.foundation/blog/2023/09/20/lf-ai-data-foundation-announces-graduation-of-openlineage-project/

[2]OpenLineage. “Object Model.” OpenLineage Documentation.
https://openlineage.io/docs/spec/object-model

[3]OpenLineage. “Facets.” OpenLineage Documentation.
https://openlineage.io/docs/spec/facets/

[4]OpenLineage. “Object Model.” OpenLineage Documentation.
https://openlineage.io/docs/spec/object-model

[5]LF AI & Data Foundation. “LF AI & Data Foundation Announces Graduation of OpenLineage Project.” LF AI & Data Foundation, September 20, 2023.
https://lfaidata.foundation/blog/2023/09/20/lf-ai-data-foundation-announces-graduation-of-openlineage-project/

Frequently asked questions

01.

What is OpenLineage?

OpenLineage is an open standard for collecting data lineage as events, governed by the LF AI & Data Foundation, which it graduated from in 2023. As a job runs, an integration reports which datasets it read, what processing it performed, and which datasets it wrote. Integrations ship for Apache Airflow, Apache Spark, Apache Flink, dbt, and a range of data warehouses. The standard defines a common format so lineage can move between tools without proprietary adapters.

02.

Is OpenLineage runtime-only?

No, and the claim circulates widely enough to be worth correcting. The specification defines job events and dataset events that carry design-time metadata, including declared inputs and outputs, schema, and source-code location. What limits most deployments is the producers rather than the standard, since integrations such as the Apache Spark listener emit when a job runs. A more accurate description is an event-emission standard for technical lineage.

03.

Why does OpenLineage exclude business glossary and governance policy?

The exclusion was deliberate and it was sound engineering. At graduation, its Technical Steering Committee members were affiliated with nine organizations including Apache Airflow, dbt, Microsoft, and Snowflake. Those organizations could agree on a schema describing what a job read and wrote, because that has a mechanical answer any engine reports the same way. They could not have agreed on what a customer record means or who owns it. A specification mandating business semantics would have died in committee.

04.

Can OpenLineage capture lineage from a mainframe?

Not through runtime capture, which is how OpenLineage is deployed almost everywhere. A COBOL batch process moving balances overnight emits nothing, since no listener runs inside it and no supported framework reports what it did. The specification does permit design-time events from static code analysis, so a tool that parses COBOL could in principle publish lineage in the format, and that remains rare in practice. Mainframe extracts, vendor file feeds, and manual reconciliation all move data that matters, and reaching them depends on analysis and modeling rather than on more runtime connectors.

05.

What does OpenLineage not provide for a BCBS 239 or DORA audit?

Three things a supervisor asks for sit outside the standard. Coverage of systems that never emit an event, since automated capture sets a ceiling. Business meaning, including ownership, definitions, and the policies attached to a field, all of which live in the institution rather than the pipeline. And the state of the lineage at a past date, which requires version control over the model and point-in-time reconstruction rather than a log of which jobs ran.

06.

How many OpenLineage events does a single pipeline generate?

A pipeline running hourly emits roughly 720 events a month, every one describing the same unchanged data flow. Across a few hundred pipelines over a year, the volume grows substantially while the number of distinct flows it describes barely moves. That behavior reflects the design working correctly, because OpenLineage records executions. Consuming platforms differ mainly in how much work they do to consolidate repeated runs into a stable model of the underlying data flow.

Published on: September 29, 2026

Contents

Related articles

Data dashboards
Blog

A sovereign cloud does not make your data sovereign

Why a residency commitment is not evidence of custody.

From Complexity to Confidence -A Modern Data Strategy in Action webinar
Blog

Your BCBS 239 evidence expires the day you file it

Three criteria that make regulatory data lineage reusable

The EU AI Act deadline has moved, but data lineage can’t wait
Blog

Your risk reports run on an unassessed supply chain

How data lineage brings assessment discipline to financial services data flows

What to look for in a data lineage platform
Blog

What to look for in a data lineage platform

A buyer's guide to choosing the foundation of your AI stack.

Data Lineage is Finally Getting the Attention
Blog

Your riskiest change ticket says low risk

Pre-change impact analysis for data estates that feed AI

Solidatus Chosen as Microsoft Purview’s Data Lineage Integration Partner
Blog

Why banks can’t meet modern regulations without data lineage

What BCBS 239 and SR 26-2 now demand at the column level

Blog

Solidatus 2026.3 puts data lineage where AI agents can reach it

MCP support, Bring Your Own LLM, and an assistant that keeps its context

Solidatus Named in 2024 FT 1000 Ranking of Europe’s Fastest-growing Companies
Blog

Lineage tools forget what your AI model saw

Bi-temporal data lineage is the foundation of forensic AI investigation

blogImgShadowAI
Blog

Shadow AI is a sovereignty problem

Why an AI use policy cannot enforce itself.

Ai lineage company
Blog

Your Data Governance Team Already Built Your AI Governance Foundation

How LSEG turned data lineage into a strategic asset for AI trust

Solidatus Named in 2024 FT 1000 Ranking of Europe’s Fastest-growing Companies
Blog

Proving data lineage to regulators

What regulators expect when they ask you to prove data lineage

Blog

Data lineage vs metadata management: The architecture behind AI governance

Why catalog-first tools fall short when regulators ask how your training data moved

Ai lineage platform
Blog

Model Risk Starts in the Data Supply Chain

When an AI model starts producing unexplainable results, the first instinct is to blame the model. Teams often rush to...

Solidatus Named in 2024 FT 1000
Blog

The Data Fairy is Dead

Five data lineage myths that the masterclass got right

Four AI governance questions your data catalog cannot answer
Blog

Four AI governance questions your data catalog cannot answer

What regulators are already asking about your AI, and what it takes to respond

Ai lineage company
Blog

The Engineering of Trust: Why Metadata isn’t enough for AI

Three questions your AI governance approach must answer

FT 1000 Ranking of Europe’s Fastest-growing Companies
Blog

The 48-Hour Test: Does Your AI Have Complete Data Lineage?

Three institutions receive the same question during Model Risk Management reviews: “Walk us through the complete data lineage for your...

Data lineage and AI
Blog

How data lineage prevents AI failures in financial services

Solidatus’ Tina Chace and fellow experts reveal why 90% of AI model failures trace back to upstream data changes

Blog

Why Data Lineage is Essential for AI: 7 Governance Challenges Solved by AI-Ready Lineage

AI-ready data lineage is a comprehensive, auditable record of how data flows through your organization, designed to support AI governance...

ai data lineage
Blog

Detailed Expectations Around End-to-End Data Lineage and BCBS 239 From the European Central Bank RDARR Guide

In May 2024, the ECB released its ‘Guide on effective risk data aggregation and risk reporting (RDARR)’...

Blog

Why is Advanced Data Lineage Fundamental for Financial Services Organizations?

Read why advanced data lineage is crucial for business success

Blog

Continuing Innovation in Advanced Data Lineage to Help Answer Business Questions

An update on some recent developments in our latest product releases

Blog

Unveiling the Path: Why Data Lineage is Crucial for Building Effective AI Products

Read more about data lineage and its business impact, including on AI, BCBS 239 and more

Blog

Solidatus & Microsoft Purview: Elevating Data Governance in the AI Era

Solidatus data lineage partners with Microsoft Purview to help enterprises trust their data

Solidatus Recognized in 2022 Gartner® Market Guide
Blog

Blasting Off: Why Proactive Data Governance is Propelling Innovation

Read our key takeaways from Gartner D&A Summit 2024

Solidatus Named in 2024 FT 1000 Ranking of Europe’s Fastest-growing Companies
Blog

Live Demo: Explore our All-new Interface

Video introducing our new interface and core features like Connected Catalog and Data Map

Blog

Visualize Snowflake Horizon and Enhance its Impact

Read about Solidatus and Snowflake Horizon's governance solution

What to look for in a data lineage platform
Blog

Advanced Data Lineage: The Cornerstone of Modern Data Governance

Explore the various aspects of data lineage and its crucial role in your organization.

Achieving Basel III Compliance: A 3-Step Action Plan
Blog

Achieving Basel III Compliance: A 3-Step Action Plan

Basel III is changing – are you prepared? Read 3 easy steps with Solidatus

Solidatus Named in 2024 FT 1000 Ranking of Europe’s Fastest-growing Companies
Blog

Building a Data Community

Read how we helped successfully launched the Houston Women in Data Chapter

Blog

Data Lineage for Better Planning

Exploring the parallels between urban planning and data planning projects

Blog

Data Distress: Data Leaders on Brink of Quitting Jobs

71% of senior data leaders in financial services polled are close to quitting their jobs

Blog

Supercharging your Snowflake Governance with Solidatus

Take a look at what's new in our partnership with Snowflake

Blog

The Value of Data: Reflections from Attending Gartner

VP Product, Tina Chace, reflects on the Gartner conference, covering data governance and AI

Blog

Douze Points for New Way to Visualize Eurovision Data

We’ve linked the Eurovision Song Contest to the realm of data governance and data lineage

Blog

Quick Answer: What is Active Metadata?

In the latest Gartner® research note, find out what active metadata is

Blog

The Amazing World of Active Metadata

The role of metadata, dynamic visualization and inference across metadata

Blog

5 Ways Great Metadata Connectors are Game-Changers

Automatic connectors are essential for efficiently mapping metadata but not all are created equal. We look at the most important...

Blog

5 Things we Can’t Wait to Do at the Gartner Summit, Orlando

We discuss injecting active metadata into your governance and 4 other things we’re looking forward to at the Gartner® Data...