Your BCBS 239 evidence expires the day you file it

From Complexity to Confidence -A Modern Data Strategy in Action webinar

Three criteria that make regulatory data lineage reusable

Thirteen years in, the supervisory scorecard has not moved

Banks have been working on BCBS 239 for thirteen years, and supervisors say they are no better at it than they were a year ago. The Basel Committee published the principles for risk data aggregation and risk reporting (RDARR) in January 2013, with global systemically important banks expected to comply by 2016.1 Banks have built programs against them and reported progress to their boards every year since. In November 2025, the European Central Bank recorded what all of that produced in its annual Supervisory Review and Evaluation Process (SREP).2

“The SREP 2025 outcomes point to persistent deficiencies in banks’ RDARR frameworks, with no improvement in the relevant average sub-score as compared with last year.” — European Central Bank, Supervisory priorities 2026-28

A flat sub-score looks like an underinvestment problem, fixable with more budget and another remediation program. Right? Wrong. Most large banks have already tried that, more than once, and the 2025 SREP result is what came of it.

The institutions that turned regulatory work into an asset the business uses daily did not outspend the ones still rebuilding the same evidence for every new request from regulators. The smart banks made an early decision about how the work would be structured, and that settled whether anything they built survived the supervisory examination that prompted it.

Banks rebuild the same evidence for every supervisory request

Most banks never decide how their compliance evidence should be structured. The shape of the program decides for them, because a remediation program ends when the examiner closes the finding that started it. An examiner picks a regulatory return, such as a credit risk or liquidity report, and asks where each figure came from. The program traces that reporting chain as it stood on the filing date. Mapping documents, spreadsheets, and screenshots satisfy the finding, and the team moves on. Every step is rational, and together they produce an evidence pack that answers one question once. So what happens when the next question is not the same one?

It rarely is. A different framework names different reports, a supervisor asks how a figure looked at a quarter-end two years back, or a business line reorganizes and the systems underneath it change. Evidence assembled for the original scope does not stretch that far, so a new program starts and rebuilds most of what the last one produced. RDARR remediation becomes a recurring cost.

The Basel Committee explained why this happens in a January 2026 newsletter summarizing what supervisors and banks reported during outreach sessions. It states plainly that it introduces no new expectations, so it describes what supervisors see rather than adding an obligation.3

“Legacy systems, distributed data estates and the dynamic nature of data lineage complicate banks’ efforts to confirm end-to-end data traceability.”— Basel Committee on Banking Supervision, Newsletter No. 36, January 2026

Legacy systems sit beyond the reach of any scanner. Enterprise data spreads across environments never designed to be read together, and lineage shifts underneath the documentation as fast as it is written. A platform limited to automated scanning stops at the first obstacle, so what the bank files is accurate for the systems it could reach and silent about the rest.

So what makes a lineage model answer the next question?

Building for reuse changes what the program leaves behind. Reusable data lineage is a model built so evidence assembled for one regulatory question can answer later ones without being rebuilt. A bank cannot add that afterward. It has to be there while the model is built, and one supervisory question is enough to test for it. A risk examiner asks which reports and models consume a particular counterparty rating field, and what that chain looked like on the day the figure was filed. Could your lineage answer that this week? Answering takes three things.

  1. Granularity at the level the question is asked: Supervisors ask about reported figures, and each one traces back to a column. A record of which systems feed which others can confirm that a risk platform draws on a customer database, without following one counterparty rating field through three transformations into the report using it. When the question is about a number, only a model kept at the level of the individual data field can answer it.
  2. Business meaning bound to the technical path: Regulations name business concepts, and no system stores a column called counterparty credit exposure. Someone has to know which physical fields carry it, and when that knowledge leaves with the analysts who staffed the last program, the next one begins by rediscovering it. A model that holds both lets a new framework map onto structures that already exist, so the mapping is done once rather than repeated.
  3. A record that holds its own history: The figure was filed on a date and the data environment has moved on since, with systems retired, transformations rewritten, and field definitions changed. A model that shows only today’s flows cannot tell an examiner what the chain looked like at the quarter-end under review, and rebuilding it from change tickets is the manual work the program was meant to end. A model that stores its own versions answers by rewinding to that date. Lineage kept this way is called bi-temporal, meaning it records when the data moved and when the model describing that movement changed.

Each is available on its own. Scanners produce technical detail with no business meaning attached, and catalogs hold business definitions without a trace back to source. Keeping all three in one model is what makes the second request cheaper than the first, and it is what most banks discover they skipped.

The difference shows up when supervisors come back

A bank could once miss one of the three criteria at little cost, because supervisors rarely asked about the same data environment twice. The ECB approved a system-wide strategy in December 2024 that monitors every supervised bank against the risk data expectations and tracks remediation plans. It begins with management accountability and widens into data quality management and IT architecture.4 Banks under it face supervisors already holding the last round’s findings, with targeted inspections where those were severe. It extends the intensifying supervisory approach the ECB signalled earlier.

When the evidence was assembled for a single examination, answering the follow-up means reopening that program and repeating much of the mapping, usually with a new team. A model built to be queried turns the same follow-up into a query against a model that already exists, and the cost difference between those two paths grows with every round of supervisory attention.

HSBC’s wholesale credit and lending work shows what a model built to be queried produces at scale. The team built it to demonstrate traceability from source to consumption across the lending book, plainly regulatory work. Solidatus’s account describes the starting point as a static, single use case and the outcome as a core pillar of the group’s data strategy, with the same model now supporting environmental, social, and governance reporting, liquidity calculations, and risk-weighted asset optimization.5 The regulatory requirement bought an asset the bank now uses daily, and the scale of that model is covered in our earlier piece on why banks cannot meet modern regulations without data lineage.

“With Solidatus, I can visualize the data model for each business outcome, manage the requirements of key stakeholders and provide greater clarity over the intricacies involved in addressing their requests.”— Sid Mubashar, Head of Wholesale Credit & Lending Data and Monitoring, HSBC

Mubashar’s phrase, each business outcome, carries the whole argument. A single model serves many of them, queried by people who had no part in building it, which is the practical test of whether regulatory work produced something durable.

Where to start before the next request lands

Three moves separate a lineage model that answers once from one that keeps answering. Each assumes lineage held as a versioned model rather than a folder of documents, which is what most remediation programs never produce.

  1. Diff the current model against the version you filed: Pull up the lineage as it stood at the last examination and compare it to today. The systems, fields, and transformations that changed since you filed show where the next inspection will find drift. Programs that kept only a document set have no earlier version to measure against.
  2. Branch a proposed remediation before committing to it: Model the fix as a future state alongside production, then check which reports and consumers it touches before anything moves. Remediation plans carry an implicit promise that the fix breaks nothing else, and a branch is where that gets tested rather than assumed.
  3. Attach the regulation’s own vocabulary to the fields it governs: Tag physical fields with the business concepts a supervisor uses, so the next framework maps onto structures that already exist. This decides whether the next program starts from an inventory or from a model, and it is best done while the analysts who know the mapping are still there.

Start with one domain already under supervisory attention rather than the whole data ecosystem. Prove the model answers a question you have been asked, then extend it.

The rebuild is the cost

Thirteen years of programs, and a sub-score that will not move, describe an industry paying repeatedly for the same work. Each program assembles evidence for the question in front of it, closes, and leaves the next team to start over, so effort keeps rising while the score holds flat. Boards now answer for that cost directly, since the ECB strategy starts with management accountability before it reaches data quality or architecture. The banks that escaped the cycle were not working with simpler data ecosystems. They built the first model so it could answer the second question, and the work stopped recurring as a cost.

Supervisors have said what they are looking for, and the Basel Committee has named traceability from origin to final use as what confirms data quality. The next examination is already scheduled. Will answering it mean running a query, or staffing another program? That was settled when the last model was built.

Solidatus is a data lineage platform built for regulated enterprises, used by banks including HSBC, Royal London Asset Management, and Bank of New York to model lineage at the individual data field, with version history that makes a past filing reproducible years later. To find out whether your own lineage could answer a second examination without rebuilding, request the banking use case overview.

1Basel Committee on Banking Supervision. “Principles for effective risk data aggregation and risk reporting.” Bank for International Settlements, January 2013. https://www.bis.org/publ/bcbs239.htm. Compliance expectation referenced in Basel Committee on Banking Supervision, “Progress in adopting the Principles for effective risk data aggregation and risk reporting,” November 28, 2023. https://www.bis.org/bcbs/publ/d559.htm.

2European Central Bank. “Supervisory priorities 2026-28.” ECB Banking Supervision, November 2025.
https://www.bankingsupervision.europa.eu/framework/priorities/html/ssm.supervisory_priorities202511.en.html.

3Basel Committee on Banking Supervision. “Implementation of the Principles for effective risk data aggregation and risk reporting (BCBS 239 Principles).” Newsletter No. 36. Bank for International Settlements, January 6, 2026. https://www.bis.org/publ/bcbs_nl36.htm. The newsletter states that it “is for informational purposes only and does not constitute new supervisory guidance or expectations.”

4European Central Bank. “Supervisory priorities 2026-28.” ECB Banking Supervision, November 2025.
https://www.bankingsupervision.europa.eu/framework/priorities/html/ssm.supervisory_priorities202511.en.html.

5Solidatus. “Solidatus Models HSBC’s Global Lending Book.” Case study, 2024.
https://www.solidatus.com/resource/case-studies/solidatus-models-hsbcs-global-lending-book/.

Frequently asked questions

01.

What does BCBS 239 require for data lineage?

The Basel Committee’s principles for risk data aggregation and risk reporting, published in January 2013, require banks to demonstrate that risk figures can be traced from the systems that captured them to the reports that use them. Global systemically important banks were expected to comply by 2016. The principles leave the technology open. The European Central Bank’s supervisory expectations have since pushed the standard of proof down to the level of individual data attributes, a level spreadsheets and dataset-level catalogs were never designed to reach.

02.

Why have banks made so little progress on risk data compliance?

The European Central Bank reported in November 2025 that SREP 2025 outcomes show persistent deficiencies in bank risk data frameworks, with no improvement in the relevant average sub-score year over year. The shortfall is structural rather than financial. Most remediation programs assemble evidence scoped to the single examination that prompted them, so when a new regulation, a new reporting period, or a reorganization arrives, the next program rebuilds much of the same work from the beginning.

03.

What is the difference between dataset-level and column-level data lineage?

Dataset-level lineage records which systems feed which other systems, so it can confirm that a risk platform draws on a customer database. Column-level lineage follows an individual data field through each transformation between origin and report. Supervisors ask about figures, and a figure lives in a column, so only a data lineage model kept at field level can answer the question an examiner asks. Column-level detail is what turns a system diagram into usable regulatory evidence.

04.

Can data lineage work built for compliance be reused for other purposes?

Yes, when the model is built for it. HSBC’s wholesale credit and lending model began as regulatory work demonstrating traceability from source to consumption, and Solidatus’s account of the engagement describes it moving from a static, single use case to a core pillar of the group’s data strategy, now supporting environmental, social, and governance reporting, liquidity calculations, and risk-weighted asset optimization. Reuse depends on three criteria being met at build time: field-level granularity, business meaning bound to the technical path, and stored version history.

05.

What did the Basel Committee say about data lineage in 2026?

Newsletter No. 36, published on January 6, 2026, states that data lineage, meaning the traceability of data from its origin to its final use, is important for confirming data quality. It also identifies legacy systems, distributed data estates, and the changing nature of lineage as obstacles to confirming end-to-end traceability. The newsletter states explicitly that it is for informational purposes only and does not constitute new supervisory guidance or expectations, so it reports what supervisors and banks are seeing rather than adding a requirement.

06.

Where should a bank start building reusable data lineage?

Start with one domain already under supervisory attention rather than the whole data ecosystem, and prove the model can answer a question the bank has already been asked. Three moves make the work reusable: compare the current lineage against the version filed at the last examination to find drift, branch a proposed remediation and check downstream consumers before committing to it, and tag physical fields with the business vocabulary supervisors use so the next framework maps onto structures that already exist.

Published on: September 3, 2026

Contents

Related articles

The EU AI Act deadline has moved, but data lineage can’t wait
Blog

Your risk reports run on an unassessed supply chain

How data lineage brings assessment discipline to financial services data flows

What to look for in a data lineage platform
Blog

What to look for in a data lineage platform

A buyer's guide to choosing the foundation of your AI stack.

Data Lineage is Finally Getting the Attention
Blog

Your riskiest change ticket says low risk

Pre-change impact analysis for data estates that feed AI

Blog

Why banks can’t meet modern regulations without data lineage

What BCBS 239 and SR 26-2 now demand at the column level

Blog

Solidatus 2026.3 puts data lineage where AI agents can reach it

MCP support, Bring Your Own LLM, and an assistant that keeps its context

Blog

Lineage tools forget what your AI model saw

Bi-temporal data lineage is the foundation of forensic AI investigation

blogImgShadowAI
Blog

Shadow AI is a sovereignty problem

Why an AI use policy cannot enforce itself.

Ai lineage company
Blog

Your Data Governance Team Already Built Your AI Governance Foundation

How LSEG turned data lineage into a strategic asset for AI trust

Blog

Proving data lineage to regulators

What regulators expect when they ask you to prove data lineage

Blog

Data lineage vs metadata management: The architecture behind AI governance

Why catalog-first tools fall short when regulators ask how your training data moved

Blog

Model Risk Starts in the Data Supply Chain

When an AI model starts producing unexplainable results, the first instinct is to blame the model. Teams often rush to...

Blog

The Data Fairy is Dead

Five data lineage myths that the masterclass got right

Four AI governance questions your data catalog cannot answer
Blog

Four AI governance questions your data catalog cannot answer

What regulators are already asking about your AI, and what it takes to respond

Ai lineage company
Blog

The Engineering of Trust: Why Metadata isn’t enough for AI

Three questions your AI governance approach must answer

Blog

The 48-Hour Test: Does Your AI Have Complete Data Lineage?

Three institutions receive the same question during Model Risk Management reviews: “Walk us through the complete data lineage for your...

Data lineage and AI
Blog

How data lineage prevents AI failures in financial services

Solidatus’ Tina Chace and fellow experts reveal why 90% of AI model failures trace back to upstream data changes

Blog

Why Data Lineage is Essential for AI: 7 Governance Challenges Solved by AI-Ready Lineage

AI-ready data lineage is a comprehensive, auditable record of how data flows through your organization, designed to support AI governance...

ai data lineage
Blog

Detailed Expectations Around End-to-End Data Lineage and BCBS 239 From the European Central Bank RDARR Guide

In May 2024, the ECB released its ‘Guide on effective risk data aggregation and risk reporting (RDARR)’...

Blog

Why is Advanced Data Lineage Fundamental for Financial Services Organizations?

Read why advanced data lineage is crucial for business success

Blog

Continuing Innovation in Advanced Data Lineage to Help Answer Business Questions

An update on some recent developments in our latest product releases

Blog

Unveiling the Path: Why Data Lineage is Crucial for Building Effective AI Products

Read more about data lineage and its business impact, including on AI, BCBS 239 and more

Blog

Solidatus & Microsoft Purview: Elevating Data Governance in the AI Era

Solidatus data lineage partners with Microsoft Purview to help enterprises trust their data

Blog

Blasting Off: Why Proactive Data Governance is Propelling Innovation

Read our key takeaways from Gartner D&A Summit 2024

Blog

Live Demo: Explore our All-new Interface

Video introducing our new interface and core features like Connected Catalog and Data Map

Blog

Visualize Snowflake Horizon and Enhance its Impact

Read about Solidatus and Snowflake Horizon's governance solution

What to look for in a data lineage platform
Blog

Advanced Data Lineage: The Cornerstone of Modern Data Governance

Explore the various aspects of data lineage and its crucial role in your organization.

Blog

Achieving Basel III Compliance: A 3-Step Action Plan

Basel III is changing – are you prepared? Read 3 easy steps with Solidatus

Blog

Building a Data Community

Read how we helped successfully launched the Houston Women in Data Chapter

Blog

Data Lineage for Better Planning

Exploring the parallels between urban planning and data planning projects

Blog

Data Distress: Data Leaders on Brink of Quitting Jobs

71% of senior data leaders in financial services polled are close to quitting their jobs

Blog

Supercharging your Snowflake Governance with Solidatus

Take a look at what's new in our partnership with Snowflake

Blog

The Value of Data: Reflections from Attending Gartner

VP Product, Tina Chace, reflects on the Gartner conference, covering data governance and AI

Blog

Douze Points for New Way to Visualize Eurovision Data

We’ve linked the Eurovision Song Contest to the realm of data governance and data lineage

Blog

Quick Answer: What is Active Metadata?

In the latest Gartner® research note, find out what active metadata is

Blog

The Amazing World of Active Metadata

The role of metadata, dynamic visualization and inference across metadata

Blog

5 Ways Great Metadata Connectors are Game-Changers

Automatic connectors are essential for efficiently mapping metadata but not all are created equal. We look at the most important...

Blog

5 Things we Can’t Wait to Do at the Gartner Summit, Orlando

We discuss injecting active metadata into your governance and 4 other things we’re looking forward to at the Gartner® Data...