The license agreement was written for a person looking at a screen. Your pipeline doesn’t have eyes.
An architecture review board approves a new enterprise analytics platform. The diagram shows a market data feed entering a cloud bucket, undergoing transformation, and being exposed via an internal API gateway. The design looks clean. Six months later, procurement receives an audit letter from the data vendor. The platform has no named users. It has service accounts. The license covers display use by natural persons. The vendor’s audit team counts every service account as an undeclared user.
For decades, market data licensing was a perimeter problem. You bought a terminal, put it in a room, and restricted the door. The architecture diagram and the license diagram were identical. Today, they are entirely different documents. A market-data feed can pass through a dozen systems while appearing as a single line in an architecture diagram. Every design decision that moves data away from a human looking at a screen moves you into a different licensing category.
Engineers assume that if data stays inside the corporate firewall, it is internal use. Vendors define internal use by corporate legal structure, not by network topology. Affiliates, subsidiaries, and joint ventures count as third parties. Consider a custodian bank that provides data to fund managers as part of a reporting service. The data vendor’s license prohibits redistribution to third parties. By serving that data to a separately capitalized fund manager, the custodian becomes a redistributor. The fund managers do not know their data is licensed. The custodian’s contract with them does not mention it either. The traffic never left the network, but it crossed a legal boundary.
Vendors distinguish between display use and non-display use. Display is a human looking at a screen. Non-display is a machine consuming the data for algorithmic trading, risk management, or valuation. Non-display fees are higher and audited aggressively.
Then there is derived data. A quantitative research desk builds a custom index to track commodity spreads using raw futures prices from an aggregator. They publish the index internally on a dashboard. The aggregator argues the index is a commercial substitute for their own proprietary analytics feed and demands a redistribution fee. The internal legal team spends months arguing whether the mathematical transformation was complex enough to sever the licensing tie. Math does not automatically launder market data. The word “derived” describes a calculation; the agreement determines the rights. If the original raw data can be reverse-engineered from the new metric, the vendor claims the derived work falls under their license.
Cloud and AI did not create the licensing problem. They made it legible. The data was always flowing to places the contract did not anticipate. Now the vendor can see the query volumes. Feeding historical tick data into an internal machine learning model triggers non-display usage clauses. If the model generates trading signals, it crosses into derived data. Training models on vendor data usually requires a separate commercial agreement. Deleting the original feed from a workspace does not settle what persists in embeddings, model artifacts, or backups.
Cloud infrastructure is designed to spin up resources on demand. Market data contracts are often written around per-server or per-application pricing. When a container cluster scales horizontally to run weekend backtests, the infrastructure bill goes up by a fraction. The market data licensing liability multiplies. To a software engineer, inserting a cache between a market data feed and an internal application is a standard pattern to reduce latency. To a data vendor’s auditor, that cache is an unauthorized internal redistribution node.
The difference between what the contract says and what the architecture does is the audit finding. The vendor’s audit team has a data flow diagram. Do you? When an exchange exercises its right to audit a regional pension fund, the auditor asks for a list of all applications consuming their real-time feed. The IT director provides a list of trading terminals. The auditor then asks for the network routing tables and firewall logs for the VLAN where the feed terminates. They find a daily batch job copying the raw feed to a shared storage drive accessible by the risk modeling team in another subsidiary. The exchange issues a bill for three years of backdated non-display enterprise usage.
The audit does not ask what you intended. It asks what you deployed. If no one recorded which feed populated a table, which jobs read it, and who received the outputs, a license review becomes a reconstruction exercise.
Entitlement enforcement is an architecture property, not a compliance afterthought. You cannot retrofit entitlement into a platform that was designed to share everything. If you cannot enforce per-user permissions in the platform, you either over-provision and pay for everyone, or you under-enforce and accept the audit exposure. Both are expensive.
The architecture review board has a missing seat. Procurement or vendor management is rarely at the table when designs are approved. They arrive at contract renewal with a bill and no leverage. A pipeline should reject data if the destination lacks the correct commercial entitlements. If your data lineage tool does not track commercial entitlements, you only know where the data went, not who was allowed to follow it.
