Explainability Is Not a Feature of AI in Mortgage. It Is the Condition of Using It.
Vendors selling artificial intelligence into mortgage operations tend to present explainability as a refinement: a capability that mature products add once accuracy is settled. That ordering is backwards. In a regulated lending process, a system that cannot account for its own output is not a less convenient system. It is frequently an unusable one.
The reasoning is worth setting out carefully, because it is often argued badly.
What the law actually requires, and of whom
The Equal Credit Opportunity Act and Regulation B require a creditor to give an applicant specific and accurate reasons when credit is denied. The Consumer Financial Protection Bureau has stated that this obligation is unchanged when the decision is produced by a complex algorithm: ECOA and Regulation B do not permit a creditor to use black-box underwriting technology where the consequence is that the creditor cannot supply those specific reasons. [1] The Bureau has been equally clear that a standard checkbox notice will not suffice where the listed reasons do not accurately describe the factors the model actually considered. [1]
Precision matters here, and it is where a good deal of vendor marketing overreaches. Adverse action obligations attach to credit decisions, not to post-closing quality control. A QC finding is not a denial of credit, and a vendor implying that a QC platform is subject to adverse-action notice requirements has the framework wrong.
The converse overreach is also worth naming, because an earlier version of this article came close to it. Saying that adverse-action rules do not govern a QC finding is not the same as saying quality control has no role in ECOA compliance. It plainly does: post-closing review is one of the places a lender tests whether credit decisions were made and documented consistently, whether stated reasons match the factors actually relied on, and whether outcomes differ across protected classes in ways that need explaining. The QC finding is not itself an adverse action. The credit decision it examines may well have been one.
What transfers is not the statute. It is the underlying principle, and post-closing quality control is one of the places a lender tests whether credit decisions were documented consistently. [3]
The pause is not a repeal
There is a further trap. The CFPB has withdrawn Circulars 2022-03 and 2023-03, along with a broader set of interpretive rules, policy statements and advisory opinions. [2] It would be an error to read that as a relaxation of the duty. The circulars were interpretive; the obligations sit in ECOA and Regulation B, and those have not changed. [1]
Be exact about the scope of the withdrawal as well. What the Federal Register notice establishes is that those guidance documents were withdrawn. It is not evidence of a blanket suspension of supervision or enforcement in this area, and this article previously implied more than the source supports. Verify current supervisory posture directly before relying on it in a policy document.
Institutions that treat a shift in supervisory posture as a change in the underlying law tend to discover the distinction at examination, or in litigation, some years later. A prudent lender builds to the statute, not to the current appetite for enforcing it.
Why quality control has its own explainability problem
Post-closing quality control operates under a different regime: agency selling guides, investor requirements, internal credit policy, but it produces findings that must survive exactly the same kind of scrutiny.
A QC finding is an assertion that something in a closed loan is wrong. That assertion travels. It goes to a remediation queue, into management reporting, into the file an examiner may sample, and sometimes into a dispute with an investor over whether a repurchase demand is justified. At every one of those points, someone will ask why.
If the answer is that a model flagged it, the finding is worth very little. It cannot be actioned by a reviewer who does not know what to fix. It cannot be rebutted by a lender who believes it is wrong. It cannot be produced to a third party as evidence of anything. A system that emits conclusions without the reasoning behind them has not reduced the lender's risk. It has converted a documentation problem into an argument the lender cannot win.
What a defensible finding contains
The test is reconstruction. Given a finding, can a competent reviewer who was not present rebuild how it was reached?
That requires four things, and all four:
- The documents compared. Which pages, from which file, in which version.
- The values read from each. What the system believed each document said, so a
misreading is distinguishable from a genuine inconsistency.
- The rule applied. Stated explicitly, attributable to its source: an agency
requirement, an investor guideline, a credit policy, and versioned, so a file reviewed in March can be shown to have been assessed under March's rules.
- Why the comparison failed. Not a category label, but the specific relationship that
did not hold.
A finding carrying all four is evidence. A finding carrying none is an opinion with a confidence score attached.
Where confidence scores mislead
Probabilistic output is useful for triage and dangerous as a substitute for reasoning. A score of 0.94 tells a reviewer how certain the system is. It does not tell them what the system concluded or why, and it cannot be shown to an examiner as the basis for anything.
Worse, confidence is frequently reported at the wrong granularity. A model may be highly confident it has read a figure correctly while the finding that depends on that figure rests on a rule nobody has written down. Confidence in extraction is not confidence in the conclusion. Presenting the first as though it were the second is a category error, and it is common.
The design consequence
If findings must be reconstructable, the architecture is constrained in a specific way: the rules cannot live only in learned parameters.
This does not mean abandoning machine learning. It means locating it correctly. Models are well suited to reading documents: classifying them, locating fields, handling the variation of real-world files. That is a perception problem and models are good at perception problems.
The adjudication that follows should be explicit. Which values must agree, within what tolerance, under which policy, is a matter of stated rules that a compliance officer can read, challenge and version. Systems built this way can use the most capable models available without inheriting their opacity, because the model's output is an input to the decision rather than the decision itself.
The position this leads to
Explainability is often framed as a tax on performance: the price of operating in a regulated industry. It is more accurately a design discipline, and it produces better systems.
A platform that must justify every finding cannot hide behind aggregate accuracy. Its errors are visible and therefore correctable. Its rules are legible and therefore improvable. Its output is defensible and therefore useful to the person who has to act on it.
The industry does not need mortgage AI that is merely accurate. It needs mortgage AI that can be examined: by a reviewer, by an auditor, by a regulator, and ultimately on behalf of the borrower, who has least visibility of all into whether their file was assessed correctly.
Sources
- Regulation B, 12 C.F.R. § 1002.9. Adverse-action notice provisions and the official interpretation of specific reasons. ECOA, 15 U.S.C. § 1691.
- Interpretive Rules, Policy Statements, and Advisory Opinions; Withdrawal, Federal Register, 12 May 2025: lists the withdrawal of CFPB Circulars 2022-03 and 2023-03. The withdrawal removes those guidance documents. It does not alter the statutory and regulatory obligations under ECOA and Regulation B, and is not evidence of a blanket supervisory pause.
- Fannie Mae Selling Guide, D1-3-01, Lender Post-Closing Quality Control Review Process. Requires both random and discretionary selection; discretionary reviews supplement the random sample rather than replace it.