Your mainframe is not going to emit OpenLineage events
What runtime capture reaches, and what you have to model yourself.
OpenLineage is an open standard for collecting data lineage as events. As a job runs, an integration reports which datasets it read, what processing it performed, and which datasets it wrote. It graduated from the LF AI & Data Foundation in 2023, and integrations now ship for Apache Airflow, Apache Spark, Apache Flink, dbt, and a range of data warehouses.1
Support for the standard has spread far enough across governance platforms that it no longer separates them, which makes “do you support OpenLineage” a weak question to build an evaluation around. The more useful question is what the standard was designed to do, because that determines what it will never do for you, regardless of how many connectors ship.
Enterprise metadata has produced few standards that anyone adopts, and OpenLineage is a genuine exception. Understanding why it succeeded also explains what it leaves for you to solve.
OpenLineage’s object model rests on three entities.2
As a job executes, it emits run events describing what it read and what it wrote.
Richer metadata attaches through facets, extensible JSON objects covering technical and operational detail such as schema, SQL, and column-level lineage. Organizations can define custom facets for anything the spec does not cover, and the documentation notes that “custom facets can be promoted to the standard by including them in the spec.”3 That extension path lets the standard grow without a central authority deciding what belongs.
Is OpenLineage runtime-only? The claim circulates widely, and it is wrong. Job events and dataset events carry design-time metadata, including declared inputs and outputs, schema, and source-code location.4 What limits most implementations is the producers, not the specification, since integrations like the Apache Spark listener emit when a job runs. The more accurate description of OpenLineage is an event-emission standard for technical lineage.
“Design lineage events (DatasetEvent, JobEvent) are not associated with a Run and represent design-time metadata.”
— OpenLineage, Object Model specification
A COBOL batch process moving balances overnight emits nothing, because there is no listener inside it and no supported framework reporting what it did. Every producer in wide deployment reads a modern engine, whether that is Spark, Airflow, dbt, Flink, or a cloud warehouse, and none of them reaches the mainframe. The standard does leave a second path open, since a tool that parses source code can emit design-time events describing what the code would do. That path asks you to analyze and model the code rather than instrument the runtime, which is a different discipline with a different cost, and almost nothing in the ecosystem does it today.
Inside a tool that emits, OpenLineage does something that no amount of documentation discipline achieves. It reports column-level lineage as a byproduct of the job running, without anyone maintaining a diagram. A Spark job that gains a new join produces events describing the new join the next time it executes, because the integration reads the execution plan rather than a description somebody wrote down.
That property solves the problem that defeats most manual lineage efforts. Handwritten documentation describes the pipeline as it was understood on the day someone wrote it, and pipelines change faster than the documentation about them. Lineage emitted by the engine cannot drift from what the engine did, since it comes from the same execution.
The format matters as much as the capture. Because the events are a published open specification rather than a proprietary export, lineage collected from Spark can be read by any tool that implements the standard, and the organization doesn’t have to negotiate an adapter for every pair of systems it owns. For teams standardizing on a modern pipeline stack, that portability is worth having on its own terms.
Both properties hold only within the same limit. They apply to the tools instrumented to emit, and to the flows that have already run. What follows comes from that scope rather than from any shortcoming in how the standard was built.
The specification defines no business glossary and no linkage to governance policy, and it says nothing about version control or point-in-time history. Reading that list as a set of shortcomings misreads how standards get built.
When OpenLineage graduated, its Technical Steering Committee members were affiliated with Apache Airflow, Apache Iceberg, Apache Parquet, dbt, Egeria, Marquez, Microsoft, Snowflake, and Superconductive.5 Those nine organizations could agree on a schema describing what a job read and what it wrote, because that question has a mechanical answer any engine can report the same way. They could never have agreed on what a customer record means or who owns it. They would have disagreed again on which retention policy governs it, and on how to show that it meant something different in March. Those questions have institutional answers rather than mechanical ones, and every one of the nine would have answered differently.
A specification that tried to mandate business semantics would have died in committee, and those projects would have kept their proprietary formats. Narrow technical scope is why OpenLineage reached graduation, and why a market that agrees on almost nothing was able to adopt it. The tradeoff is straightforward: the scope that made the standard adoptable is the same scope that leaves everything else to whoever runs the data environment.
A pipeline running hourly emits roughly 720 events a month, every one describing the same unchanged data flow. Run that across a few hundred pipelines for a year and the volume becomes substantial, while the number of distinct flows it describes barely moves.
None of that indicates a defect. OpenLineage records executions because executions are what it was built to record. An architect looking for a map of the enterprise data landscape is asking a different question of the same material, and converting a record of ten thousand job runs into a model of how data moves is work that somebody has to do. Platforms that consume the standard differ mainly in how much of that conversion they perform for you, and how repeated runs get consolidated into one versioned flow is a fair thing to ask any of them to demonstrate.
Granularity compounds the conversion problem. The specification permits flexibility in what a producer includes, so some integrations provide column-level detail while others provide only table-level detail. Coverage across your pipelines ends up uneven in ways that depend on which framework emitted which event, and whatever consumes the stream has to reconcile that unevenness.
The sharper constraint shows up before deployment. A flow that has not run yet has emitted nothing, so a proposed schema change has no event history to analyze against. Teams that want to know what a change will break have to model the intended flow rather than wait for it to execute and read the wreckage afterward, which is the difference between impact analysis and incident review.
Supervisors under BCBS 239 and the Digital Operational Resilience Act (DORA) ask questions that a well-instrumented event stream cannot answer, and higher coverage percentages don’t resolve them.
Which systems feed this regulatory report, including the ones that never emit an event? Mainframe extracts, vendor feeds arriving as files, and the manual reconciliation a team performs at quarter end all move data that matters, and none of them appear in any event stream. Automated capture sets a ceiling, and modeling the rest is what gets you past it.
What does this data element mean, and who is accountable for it? Ownership, business definitions, and the policies attached to a field live in the institution rather than in the pipeline, and no facet emitted by Spark supplies them.
The hardest question concerns time. What did the lineage look like in March, before the schema change, when the report was filed? Answering requires the data lineage model as it stood on that date, carrying the governance context that applied then. A runtime log records which jobs ran in March, which tells you considerably less than how the data was governed then. What regulators expect is the state of the model, not the execution history. Version control over the lineage model and point-in-time reconstruction are what make that answerable, and neither belongs in an event-emission standard.
Map what can emit before you evaluate anything. List which systems in your data environment produce OpenLineage events at all, and treat that list as the permanent ceiling on what automated capture will ever cover for you. Everything outside it needs a modeled representation, and buying more connectors will not shrink the list.
Test a past date rather than a current diagram. Pick a regulatory report, pick a date before your last significant schema change, and try to reconstruct what fed it and who owned each element. Whether you can answer tells you whether you hold a governed data lineage model or a well-organized log.
Stitch across one system that does not emit. Take a single Spark flow, connect it upstream to the non-emitting source that feeds it, and connect it downstream to the report it reaches. That exercise shows how much of your actual data ecosystem sits outside the instrumented middle, and it costs an afternoon rather than a procurement cycle.
Solidatus ingests OpenLineage events and turns them into the governed model those three tests ask for, harmonizing repeated runs into a versioned data lineage model, stitching that model to systems the standard never instruments, and keeping the point-in-time history an examiner asks for. The OpenLineage integration page covers how the connector works.
[1]LF AI & Data Foundation. “LF AI & Data Foundation Announces Graduation of OpenLineage Project.” LF AI & Data Foundation, September 20, 2023.
https://lfaidata.foundation/blog/2023/09/20/lf-ai-data-foundation-announces-graduation-of-openlineage-project/
[2]OpenLineage. “Object Model.” OpenLineage Documentation.
https://openlineage.io/docs/spec/object-model
[3]OpenLineage. “Facets.” OpenLineage Documentation.
https://openlineage.io/docs/spec/facets/
[4]OpenLineage. “Object Model.” OpenLineage Documentation.
https://openlineage.io/docs/spec/object-model
[5]LF AI & Data Foundation. “LF AI & Data Foundation Announces Graduation of OpenLineage Project.” LF AI & Data Foundation, September 20, 2023.
https://lfaidata.foundation/blog/2023/09/20/lf-ai-data-foundation-announces-graduation-of-openlineage-project/
01.
OpenLineage is an open standard for collecting data lineage as events, governed by the LF AI & Data Foundation, which it graduated from in 2023. As a job runs, an integration reports which datasets it read, what processing it performed, and which datasets it wrote. Integrations ship for Apache Airflow, Apache Spark, Apache Flink, dbt, and a range of data warehouses. The standard defines a common format so lineage can move between tools without proprietary adapters.
02.
No, and the claim circulates widely enough to be worth correcting. The specification defines job events and dataset events that carry design-time metadata, including declared inputs and outputs, schema, and source-code location. What limits most deployments is the producers rather than the standard, since integrations such as the Apache Spark listener emit when a job runs. A more accurate description is an event-emission standard for technical lineage.
03.
The exclusion was deliberate and it was sound engineering. At graduation, its Technical Steering Committee members were affiliated with nine organizations including Apache Airflow, dbt, Microsoft, and Snowflake. Those organizations could agree on a schema describing what a job read and wrote, because that has a mechanical answer any engine reports the same way. They could not have agreed on what a customer record means or who owns it. A specification mandating business semantics would have died in committee.
04.
Not through runtime capture, which is how OpenLineage is deployed almost everywhere. A COBOL batch process moving balances overnight emits nothing, since no listener runs inside it and no supported framework reports what it did. The specification does permit design-time events from static code analysis, so a tool that parses COBOL could in principle publish lineage in the format, and that remains rare in practice. Mainframe extracts, vendor file feeds, and manual reconciliation all move data that matters, and reaching them depends on analysis and modeling rather than on more runtime connectors.
05.
Three things a supervisor asks for sit outside the standard. Coverage of systems that never emit an event, since automated capture sets a ceiling. Business meaning, including ownership, definitions, and the policies attached to a field, all of which live in the institution rather than the pipeline. And the state of the lineage at a past date, which requires version control over the model and point-in-time reconstruction rather than a log of which jobs ran.
06.
A pipeline running hourly emits roughly 720 events a month, every one describing the same unchanged data flow. Across a few hundred pipelines over a year, the volume grows substantially while the number of distinct flows it describes barely moves. That behavior reflects the design working correctly, because OpenLineage records executions. Consuming platforms differ mainly in how much work they do to consolidate repeated runs into a stable model of the underlying data flow.
Published on: September 29, 2026