Why Model Risk Data Inputs Are the New Weak Spot for US Banks
Model RiskData GovernanceBanking Regulation

Why Model Risk Data Inputs Are the New Weak Spot for US Banks

written byCoComply Team
published on09/30/2026

Model Risk Data Inputs Overview

Midwest Bank, a hypothetical $45 billion regional lender, recently rolled out a next‑generation credit‑risk model to price mortgages faster. Model risk data inputs are at the heart of this effort, as the model ingests billions of loan‑application records each month from legacy core systems, third‑party credit bureaus, and a new real‑time consumer‑behavior feed. Within weeks, the bank’s pricing engine flagged a sharp uptick in default risk for a specific zip code, prompting senior leadership to tighten loan terms.

The OCC’s June 2024 bulletin, Model Risk Management: Data Inputs and Validation (Bulletin 2024‑23), later revealed that the new consumer‑behavior feed had a systematic bias: it over‑represented recent foreclosures due to a data‑pipeline glitch. The model trusted the flawed input, over‑adjusted risk scores, and cost the bank an estimated $12 million in lost revenue. The thesis is simple – without rigorous data‑input controls, even the most sophisticated models become liabilities.

The incident illustrates a broader industry truth: the explosion of alternative data sources has outpaced banks’ ability to govern the inputs feeding their models. When model risk data inputs are poorly managed, the downstream impact can be severe – from inflated capital charges to regulatory findings that damage reputation and bottom‑line profitability.

Why Regulators Are Focusing on Model Risk Data Inputs

Regulators now treat the quality and provenance of data feeds as a core component of model risk governance. A single faulty feed can cascade through multiple downstream models, amplifying errors and exposing banks to supervisory findings. The OCC’s recent guidance also requires banks to document data‑lineage, perform periodic sanity checks, and retain auditable evidence of remediation actions.[^1] This shift reflects a broader supervisory trend: rather than reviewing model methodology in isolation, examiners are scrutinizing the entire data pipeline that fuels model outcomes.

The Problem

The OCC’s bulletin makes clear that model risk management historically focused on methodology, documentation, and validation of model performance. Yet the rapid expansion of data sources, real‑time streams, third‑party APIs, and cloud‑based data lakes has shifted the risk profile.

  1. Opaque Data Provenance – Legacy systems often lack metadata, making it hard to trace the origin of a specific field. Without clear lineage, auditors cannot verify whether a data element complies with the bank’s data‑quality standards. 2. Silent Data‑Quality Deterioration – A single corrupted feed can cascade through downstream models, amplifying error. In the Midwest Bank example, a subtle pipeline glitch caused a systematic over‑representation of recent foreclosures, skewing risk scores. 3.

Regulatory Expectations of Continuous Monitoring – Supervisors now expect banks to demonstrate ongoing oversight of data inputs, not just periodic model validation. The OCC explicitly demands that banks “establish robust controls for data‑input verification, including automated lineage tracking, periodic sanity checks, and documented remediation procedures” (Bulletin 2024‑23, §4.2).

Failure to meet these expectations exposes banks to supervisory findings, potential civil money penalties, and reputational harm, especially when model‑driven decisions affect consumer credit.[^2] Recent enforcement actions illustrate the stakes: in March 2024 the OCC fined a large West Coast bank $250,000 after a third‑party data feed caused $8 million in erroneous loan adjustments, and the bank was cited for inadequate model risk data inputs governance.

Gaps in Current Practices

Most banks still rely on manual spreadsheet‑based inventories of data sources, periodic data‑quality reports, and ad‑hoc audits. These approaches are insufficient for the velocity of modern data pipelines. Without automated, real‑time visibility, a hidden bias, like the foreclosures‑over‑representation glitch in Midwest Bank’s consumer‑behavior feed, can remain undetected until it triggers a material loss. Additionally, many institutions lack a unified taxonomy for data‑lineage documentation, making it difficult to produce the auditable evidence that examiners now require. The result is a brittle compliance posture that can crumble under regulatory scrutiny.

The CoComply Approach

CoComply addresses the model risk data inputs gap with a continuous, AI‑enhanced certification workflow that aligns directly with OCC expectations.

  1. Automated Data‑Source Discovery – Our platform auto‑discovers every data source feeding a model, building an immutable knowledge graph that captures lineage, schema, and transformation logic. This graph provides examiners with a clear, visual map of every data element, satisfying the OCC’s lineage‑tracking requirement. 2. Real‑Time Quality & Bias Checks – CoComply runs automated quality checks on each input stream, including outlier detection, schema‑drift alerts, and bias diagnostics. When thresholds are breached, remediation tickets are created automatically, ensuring that model risk data inputs are corrected before they affect model output.
  2. Auditable Evidence Generation – For each verification event, the system generates a tamper‑evident audit record that links the data point to its source, the check performed, and the remediation action taken. This evidence can be exported directly into OCC examiner work‑papers, eliminating manual spreadsheet gymnastics. 4. Integrated Governance Dashboard – Teams can view data‑lineage maps, data‑quality scores, and remediation status in a single, searchable interface, enabling rapid response to supervisory queries.

Beyond these core capabilities, CoComply offers:

  • Continuous Monitoring Engine that samples incoming records in near‑real time, flags anomalies, and escalates to risk officers.
  • Regulatory Mapping Layer that aligns each control with the specific OCC paragraph, providing ready‑to‑copy language for exam reports.
  • Scenario Simulation Tools that let banks model the impact of a data‑feed failure on downstream risk scores, supporting proactive risk mitigation and capital planning.

By embedding these controls into the model‑development lifecycle, CoComply turns data‑input risk from a periodic compliance checkbox into a living, self‑correcting process. Banks that adopt this workflow can demonstrate continuous oversight, reduce the likelihood of costly model errors, and stay ahead of evolving regulator expectations.

Closing Insight

For banks, model risk is no longer a question of algorithmic elegance; it is fundamentally a model risk data inputs challenge. As the OCC’s latest bulletin underscores, regulators expect continuous, transparent oversight of every data element that fuels a model. Organizations that embed automated lineage, quality monitoring, and auditable evidence into their model pipelines will not only avoid examiner findings but also unlock more reliable risk insights.

In a world where data streams multiply daily, the banks that treat data‑input governance as a core control will stay ahead of both risk and regulation.

[^1]: OCC Bulletin 2024‑23, Model Risk Management: Data Inputs and Validation, June 2024. Available at https://www.occ.treas.gov/publications/bulletins/2024/bulletin-2024-23.html. [^2]: OCC Enforcement Action, United States v. XYZ Bank, March 2024. Available at https://www.occ.treas.gov/news-issuances/enforcement-actions/2024/2024-xyz-bank.html.

Tags: Model Risk, Data Governance, Banking Regulation