The model inventory arrived on a Tuesday. It contained three hundred and forty entries. The IT architecture team recognized twelve of them. The rest were spreadsheets on shared drives, Access databases from a decade ago, and vendor analytics modules embedded inside custody platforms. The risk function had sent a questionnaire asking for every quantitative method used in business decisions. The business answered literally.
Supervisory frameworks define a model as any system that takes inputs, applies quantitative methods, and generates outputs to drive a decision. That definition was written for pricing engines and capital reserve calculations. It also catches the spreadsheet that calculates a client’s fee rebate, the Python script scoring annuity applications, and the vendor tool that adjusts month-end valuations. The framework does not distinguish between a tool built by a quantitative developer and a workbook built by a controller with a deadline.
The first model inventory in a large institution is an archaeology project. It turns up calculations the organization has run for years without a coherent owner. A central application register will list the portfolio management platform and the data warehouse. It will not list the user-built tool that takes extracts from both and produces the figures discussed at the valuation committee. When the risk function maps data lineage and infrastructure dependencies, they hand the list to IT. IT suddenly finds itself providing access controls, disaster recovery, and change logs for systems it never built, never funded, and was never told existed.
Ownership is harder to assign than custody. The business built the spreadsheet. IT runs the file share it lives on. The risk function wants a validation report and a named owner. The vendor supplied the analytics engine and will not disclose the methodology. Every party has a reasonable claim that someone else should own the model. The steering committee usually settles the dispute by assigning ownership to whoever has the least political capital.
Vendor models present a specific problem. You cannot validate what you cannot see, but the supervisor still expects validation. A steering committee meets to approve a new credit-scoring API. The vendor explains their machine learning model is a black box by design to protect intellectual property. The head of model risk points to the supervisory guidance stating the institution is ultimately responsible for any model it uses. Procurement asks if they can add a liability clause to the contract.
For these black boxes, validation focuses on observables. The risk team tests inputs, outputs, sensitivity, and stability over time. They review the vendor’s governance and the contract’s change notification clauses. The validation report becomes a document detailing the limits of the institution’s own knowledge. If the contract lacks a right to audit, change notification requirements, or escrow agreements for source code, the validation report flags the omission. Procurement becomes part of model governance when a contract renewal determines what the institution is allowed to inspect.
Traditional model validation assumes a model is built, validated, and then changed only through a controlled process. The reality of end-user computing is different. A treasury analyst’s Excel workbook with fourteen tabs and VBA macros calculates counterparty exposure. It has been running since 2016. The person who built it left three years ago. The workbook is password-protected, and the password is on a note inside a desk drawer. The current desk head pastes in a CSV extract from the trading system, hits F9, and emails the output to the regulator.
Change control for that spreadsheet is a person opening the file. Change control for a machine learning model is a pipeline that retrains itself on new data every week. The model risk framework requires re-validation for material changes. The definition of material is where the arguments happen. The framework assumes both the spreadsheet and the pipeline go through the same change advisory process. They do not.
Machine learning and generative AI make an old boundary dispute visible. Teams already argued about whether a spreadsheet was a model or a simple calculator. Now they argue whether an AI tool that summarizes credit files or proposes classifications for a review queue is a model. A business team deploys an assistant that reads manager reports and suggests triage categories. Staff approve each classification. Six months later, the queue has grown, reviewers mostly accept the proposals without reading the source documents, and no one can show whether the assistant still handles a new report format reliably.
The validation team, trained in stochastic calculus and deterministic math, asks for the statistical bounds of the output and the drift metrics. The framework was written for systems that give the same answer on Tuesday as they did on Monday. Now the risk committee is asking what validation means for a model that updates its own weights. Testing has to account for the actual task, the data encountered in production, and whether the human reviewers can actually recognize and correct the errors they are expected to catch.
An institution cannot validate a marketing script with the same rigor as a derivatives pricing engine. The inventory must be tiered based on financial impact, regulatory consequence, and mathematical complexity. The tier dictates validation frequency and documentation depth. The business wants its tools in the lowest tier to avoid the paperwork. The risk function wants them higher. If the institution applies the highest tier requirements to every tool in the inventory, the validation queue will back up for two years. The business will respond by hiding their models, which increases the actual risk.
The specific models will change. The spreadsheets will be replaced by Python scripts, which will be replaced by vendor APIs, which will be replaced by whatever comes next in machine learning. The governance structure—the model risk committee, the tiering criteria, the requirement to trace an output back to a decision—is the durable part. IT provides the infrastructure evidence, including access controls, data lineage, and change logs. The risk function provides the model evidence, including validation, performance, and limitations. Both teams have to map their evidence to the same inventory.
