The request arrives as a list of figures. Not a data request. A list of numbers already filed, now in question. A supervisory authority points to a specific aggregated exposure metric on a quarterly liquidity report and asks for the line-item trades that make up that number. The reporting team can produce the final workbook and the warehouse query. But the query now returns a different result. The difference comes from a security classification updated after filing. Nobody kept the earlier classification alongside the report run. The number was correct when the team filed it. Six weeks later, the institution cannot prove it.
Enterprise architecture diagrams show data flowing smoothly from trading platforms and policy administration systems into pristine, governed warehouses. In reality, the governed warehouse feeds a CSV extract, which feeds a desktop spreadsheet. An analyst applies a manual haircut based on late-breaking market news before uploading the result to the regulator’s portal. The pipeline is traceable. The last mile is a black box.
Every manual adjustment to a reported figure lives somewhere—a spreadsheet, an email, a note in a ticketing system. It is almost never in the lineage graph. An analyst opens a spreadsheet named for the quarter, the report, and a version number that ends in “final.” They need to explain why a figure was overridden. The override was made by someone who left the institution two years ago. The note in the cell says “check with J.”
The adjustment layer is the shadow system of record. These overrides often hold legitimate policy interpretations, not just workarounds. Replacing them with a pipeline without recording who approved the interpretation makes the process faster but not more defensible. The adjustment is part of the calculation whether it lives in code, a table, or a spreadsheet cell.
Organizations try to buy compliance by deploying automated lineage scanners across their data lakes. These tools map database schemas and ETL pipelines. They produce a complex web of nodes and edges. They also fail during a regulatory audit. The actual regulatory logic does not live in the data lake. It lives in tactical SQL scripts and manual offline adjustments. You cannot scrape lineage out of a risk manager’s head. A vendor demo shows perfect lineage from source to report using clean data with no adjustments. The institution’s data has fourteen manual adjustments. The vendor’s tool shows none of them. Catalogs map the plumbing; they do not map the human interventions.
Traceability is the evidence you gather. Reproducibility is whether you can run the same report again with the same inputs and get the same output. Most institutions can trace a number. Fewer can reproduce it six months later.
Regulators ask questions about a report submitted in the past. The source systems have since processed amendments, cancellations, and retroactive corrections. Source systems manage state. When a trade is amended on Thursday, the Wednesday record is updated. If the architecture only stores the current state, recreating the exact dataset used for that historical submission is impossible.
A custodian sends a corrected positions file after the close. Operations loads it into the same location as the first file, overwriting it. The submitted figure can be found in an emailed attachment, but the team cannot rerun the calculation against the original input without recovering it from a backup. Without a bitemporal data model—recording both when an event happened in the real world and when it was recorded in the database—lineage breaks the moment data changes retroactively.
Semantic drift compounds the problem. Data lineage is not just about moving a field from one table to another. A core banking system classifies a counterparty by internal credit rating. The risk data mart translates that into a standardized regional rating. The regulatory reporting layer groups it into a specific supervisory risk bucket. A lineage tool shows the pipe, but not the business logic that changed the meaning of the data at each hop.
Two teams report exposures using the same warehouse table. One applies the legal-entity hierarchy effective at the reporting date; the other uses the hierarchy current when the extract runs. Their totals differ. Both teams initially call the discrepancy a data-quality issue. It is a lineage failure.
Lineage is a design constraint, not a documentation exercise. You can document a pipeline after the fact. You cannot document a path that was never captured. If the transformation logic is in code, the code is the lineage. If it is in a spreadsheet, the lineage is in someone’s head.
To make a report reproducible, it needs an identity. A report run manifest must point to a frozen snapshot of inputs, code versions, reference-data versions, parameters, adjustment records, and environment. A timestamp on the final PDF is not enough. If any of those elements move, the output moves.
This means treating manual adjustments as controlled data. Every manual change to a reported figure must be recorded as an event: actor, timestamp, reason, before and after values, and approval. The adjustment becomes a first-class data object, not an email. Banning spreadsheets is a less useful requirement than banning untraceable changes.
It also requires explicit handoffs between source-system owners and reporting teams. Source owners can attest to what they supplied; reporting owners can attest to transformations and filing decisions. Neither should be asked to certify a calculation they cannot inspect. When a source system changes a field’s meaning or precision, a reverse lineage query catches the impact before it breaks a filing. Forward lineage answers the regulator. Reverse lineage prevents the problem.
A lineage map that drifts from the pipeline creates false confidence by claiming a path is known when the actual execution says otherwise, leaving the regulator to discover the difference. The cost of capturing that path is paid at design time, while the cost of failing to capture it is paid at review time.
