INSIGHTS
Article

Defining Data Risk

A second-line discipline: what it oversees, and what it leaves to the first line

Ask five banks where Data Risk’s remit begins and ends, and you will likely get five different answers, some of which contradict each other inside the same institution. That inconsistency is not a staffing problem or a maturity problem; it is a definitional one. Until the boundary between running the data and overseeing it is drawn clearly, every downstream argument about scope, headcount, and authority is really just this same unresolved question wearing a different hat.

Data Risk suffers from a definition problem, and most of the confusion traces to one unresolved question: is it an operating function or an oversight function? Our position: Data Risk is a second-line discipline that oversees, but does not run, the data. The first line (the CDO or data office) owns and operates: it classifies data, builds lineage, runs retention and privacy processes, and designs and executes controls. Data Risk sets the standards, independently challenges whether they are being met, assesses residual exposure, and reports it. Once that boundary is clear, most scope arguments dissolve.

The components

Four components make up the discipline, and they build on one another in a fairly strict order. Classification comes first because everything else depends on it; retention and privacy each draw on it in different ways; and quality, in a banking context, is where the stakes are highest. Taking them in turn shows why the sequence matters as much as the substance.

Data classification is foundational, not one component among equals; you cannot dispose of data defensibly or honor a privacy right without first knowing what you hold and how sensitive it is. The first line labels the data; Data Risk defines what an adequate classification scheme looks like and challenges whether it’s applied correctly. How access is then gated against those labels (attribute-based or role-based) is a separate first-line enforcement decision, keeping those decisions separate so classification isn’t confused with access control.

Data retention is best framed as defensible disposal: keeping what must be kept, defensibly removing the rest. The first line runs the disposal; Data Risk sets the retention standard and challenges whether disposal is actually defensible. Retention is a direct consumer of classification.

Data privacy is the second consumer of classification. The first line operates privacy processes and rights handling; Data Risk sets the expectations and challenges adherence.

Data quality is the component most often missing from these conversations and, in a banking context, arguably the core. Under BCBS 239, the weight of data risk falls on whether the data feeding risk reports is accurate, complete, and timely. The first line builds lineage and operates quality controls; Data Risk defines what adequate lineage and control coverage mean, challenges completeness and effectiveness, and assesses residual exposure against appetite. One boundary here is unsettled across firms: whether the second line independently tests key controls or only oversees and challenges the first line’s own testing, leaving independent testing to internal audit.

The prioritization spine

Risk-report materiality is not a fifth silo; it is the lens that decides where the above oversight concentrates. Data Risk works backward from the material risk reports to the critical data elements that feed them, then concentrates its classification expectations, control challenge, and lineage scrutiny on those elements. Critical data elements are the mechanism: report materiality determines which elements are critical, and that is where second-line oversight concentrates.

Measuring the residual

Oversight only means something if the residual exposure is quantified. Data Risk expresses appetite through KRIs and thresholds tied to the critical data elements behind material reports: completeness, accuracy, and timeliness measures, control-effectiveness rates, aging of open issues, and reports exposure against those thresholds to risk committees. This is what makes the model operational rather than conceptual: report materiality decides what to measure, and the metrics show whether exposure sits within appetite.

When the lines disagree

Challenge is only effective if disagreement has a defined resolution path. When Data Risk challenges a control or a classification and the first line dissents, the difference escalates through defined governance, to a data risk committee and ultimately to the accountable executive or board-level risk committee, rather than being settled informally between the lines. The second line raises the challenge; governance makes the decision.

When the boundary between execution and oversight is clear, Data Risk becomes what it should be: a function that strengthens institutional trust in data, reduces the cost of regulatory challenge, and gives leadership a reliable view of where exposure actually sits.

From the Arrayo Insights Desk