E3
Documents and evidence

Mortgage Quality Control Is a Validation Problem, Not an Extraction Problem

Shailesh Bhujbal·6 min read·Published 7 September 2026·Last reviewed 8 September 2026


The mortgage industry has invested heavily in document extraction. Optical character recognition, and more recently large language models, have made it practical to read data off a loan file at a speed and cost that would have been implausible a decade ago. That capability is real and worth having. It is also routinely mistaken for the thing it is not.

FIGURE 1Extraction is not quality controlTwo different questions, with different failure modes.EXTRACTIONWhat does this document say?Reads one document at a timeReports a confidence in the readingExcellent, and largely a commodityConfidence describes how clearly a value was read.VALIDATIONIs this file sound, under therules that apply to it?Compares values across documentsDecides from a stored, versioned ruleReproducible months laterIt says nothing about whether the value is correct in context.necessarynot sufficientA system can be excellent at the left and structurally incapable of the right. From a demo, the difference is invisible.The difference shows up later, in what escapes.
Figure 1 Extraction is not quality control

Extraction answers the question what does this document say. Quality control asks a harder one: is what these documents say true, complete, and mutually consistent. The distance between those questions is where the industry's losses accumulate, and no improvement in reading accuracy closes it.

The defect data makes the point more forcefully than the argument does.

What the industry's own defect data shows

ACES Quality Management publishes a quarterly analysis of post-closing quality control findings drawn from loans reviewed in its platform. Its Q4 and calendar-year 2025 report puts the overall critical defect rate at 1.38% for Q4, down from 1.79% in Q3, with the calendar-year average at 1.50%: essentially flat against 1.52% in 2024. [1]

The headline rate is not the interesting part. The composition is.

Defect category, Q4 2025Share of critical defects
Legal, Regulatory and Compliance24.66%
Income and Employment21.52%

Legal, Regulatory and Compliance returned to the top position, rising roughly 30% from 18.97% in the prior quarter: its third consecutive quarterly increase. Income and Employment fell from the top spot, declining about 21%. [1] Across the full year, the sharpest movements were in Borrower and Mortgage Eligibility, up 291.58% year over year, and Credit, up 166.13%. [1]

Consider what each of those categories actually requires a reviewer to do.

A compliance defect is almost never legible from one document. It is a timing relationship between an application date and a disclosure delivery date, or a fee that must be tested against a tolerance established elsewhere in the file. An income defect is a comparison between what the application asserts and what pay stubs, tax returns and a verification of employment independently support. An eligibility defect is the loan measured against a guideline that does not appear in the file at all.

None of these are reading problems. Every one of them is a comparison problem.

The defects that cost money are the cross-document ones

The repurchase data points the same direction. Analysis of GSE repurchase activity found that income-related and appraisal-related issues together accounted for 57% of repurchase demands across the eighteen months from April 2023 to October 2024, with an average cost to the seller of an estimated $32,288 per demand, a demand, not a completed repurchase, against a request rate of 0.49%. [2]

Income and appraisal are both categories where cross-document comparison is usually how the problem is found: an income question is normally resolved by holding the application against the corroborating evidence, an appraisal question by holding the valuation against the subject property, the contract and the comparables.

Two limits on that inference, stated plainly. First, a category label does not establish the mechanism: an income defect can equally be a single miscalculated figure, or a document that was never obtained, neither of which a cross-document comparison would catch. Second, the 57% is a share of demands in one study's population; it does not measure how many were avoidable, and it does not imply that any particular system would prevent that proportion. What the figure supports is where to look first. It does not support a savings estimate, and this article does not make one.

It also does not diminish extraction. Cross-document validation runs on extracted values, so extraction accuracy is a precondition for it, not an alternative to it.

Aggregate repurchase volumes have fallen substantially from their 2022 peak, Freddie Mac's total repurchase dollar volume dropped 54% between the second and fourth quarters of 2023, from $594 million to $276 million, while Fannie Mae's declined 21% over the same period, from $444 million to $349 million. [4] That decline is welcome, and it partly reflects better origination-era credit. It does not change the composition of what remains.

Why extraction coverage is the wrong metric

A vendor reporting that a platform captures several hundred fields per file is describing throughput. It is not describing assurance. The relevant questions are narrower:

  • Which of those fields is tested against an independent source?
  • What rule governs each test, and where is that rule written down?
  • When a test fails, what evidence is produced alongside the finding?
  • Can a reviewer, or an examiner, reconstruct why the system reached its conclusion?

A platform that extracts fifty fields and validates each against corroborating documents is performing quality control. A platform that extracts five hundred and validates none is performing data entry at scale.

Validation requires codified, versioned rules

If quality control is the testing of assertions, the tests must exist somewhere explicit. This is not an implementation detail. It determines whether a finding can be defended.

Rules should be versioned, so a file reviewed in March can be shown to have been assessed under the rules then in force. They should be attributable to their source: an investor guideline, an agency requirement, an internal credit policy. And they should be legible to the people accountable for them, which generally means they cannot live only inside a model's weights.

A finding that cannot be traced to a stated rule is a judgement without its workings attached. Experienced reviewers form sound judgements and can explain them; what a stated rule adds is that the explanation survives the reviewer's absence, applies identically across a team, and can be produced at volume when an examiner asks for a hundred of them.

Explainability as a precondition, not a feature

Discussions of explainable systems often treat transparency as an enhancement. In mortgage quality control it is closer to a precondition.

When a reviewer is told a file contains a defect, they must be able to see which documents were compared, which values were read from each, which rule was applied, and why the comparison failed. Without that chain the finding cannot be actioned, cannot be rebutted, and cannot be produced to an examiner. A system that returns conclusions without evidence transfers risk to the lender rather than reducing it.

That the leading defect category is now Legal, Regulatory and Compliance makes this concrete. These are precisely the findings most likely to be examined by a third party, and least defensible on assertion alone.

What follows

The trajectory of the technology is not in doubt. Extraction will keep improving and its cost will keep falling, which is why extraction is unlikely to remain a source of competitive advantage for long.

The durable question is what a system does after the reading is finished: whether it holds documents against one another, whether it applies rules a human can inspect, and whether it produces evidence that survives scrutiny. Quality control was always a validation discipline. Capable extraction has not changed that. It has removed the excuse that the comparison was too laborious to perform.

Sources

  1. ACES Quality Management, Mortgage QC Industry Trends Report, Q4 and CY 2025. Critical defect rate 1.50% for CY 2025 against 1.52% for CY 2024. Category figures are shares of critical defects, not defect incidence, and are drawn from the report's sample of reviewed loans.
  2. STRATMOR Group, Unpacking the drivers and costs of GSE repurchase demands, reported by National Mortgage News. The $32,288 figure is a study estimate of cost per repurchase demand, not per completed repurchase; income and appraisal together account for 57% of demands in that study. A category share does not establish that those demands were preventable.
  3. Fannie Mae Selling Guide, D1-3-01, Lender Post-Closing Quality Control Review Process. Requires both random and discretionary selection; discretionary reviews supplement the random sample rather than replace it.
  4. Quarterly repurchase dollar volumes reported by the GSEs and summarised in trade coverage of the 2023 repurchase cycle. We have not been able to trace these particular quarter-on-quarter figures to a primary publication we can cite, so treat the direction as well attested and the exact amounts as unconfirmed. They are not load-bearing for the argument: the point is that aggregate volumes fell while the composition of what remained did not change.

See a finding taken apart.

See how each value was read, which rule was applied, and what the record looks like later.