Leading AI Models Show Bias in Mortgage Underwriting Tests
New benchmark reveals accuracy gaps and discrimination risks as lenders rush to automate loan origination.

Three leading general-purpose AI models demonstrated significant bias when evaluating mortgage applications, flagging deposits from non-English names as potentially foreign 77% of the time compared to just 13.3% for English names, according to new research from Columbia University.
The findings arrive as more than 80% of mortgage lenders evaluate AI technology and 17% have already deployed it in production workflows, creating urgent questions about accuracy and fairness in automated lending decisions.
A new standard for testing mortgage AI
Researchers introduced MortarBench, an open-source benchmark designed to measure AI performance on common mortgage origination tasks. The tool tests whether models can correctly match payroll deposits to listed employers, identify deposits requiring scrutiny, and determine joint account ownership.
Even on these routine tasks, the strongest models struggled. Gemini 3.1 Pro achieved 77.1% accuracy on complete answers, GPT-5.5 reached 76.8%, and Claude Sonnet 4.6 scored 51.4%. Matthew Toles, a Columbia doctoral student and study co-author, emphasized these scores reflect "naive use" equivalent to pasting documents into ChatGPT—not the specialized proprietary models major lenders typically develop.
Where general AI models fail
The most revealing weakness appeared in transaction identification. Models frequently misclassified entries: labeling personal loans as buy-now-pay-later transactions, assuming all wire transfers were international, or treating one-time housing payments as recurring expenses.
"You have so many transactions; it's very easy to miss one or two," said Zhou Yu, a Columbia associate professor and co-author. But models made the opposite error too, incorrectly flagging transactions that didn't meet the criteria.
The name-origin bias represents a more troubling pattern. "That's not how America works," Toles noted. "You aren't determined whether you're foreign based on how your name sounds."
Why it matters
Mortgage lenders face competing pressures: originating a retail loan costs approximately $11,800, while basic digital underwriting can save $1,700 per loan and cut five days from production time. But speed means nothing if the work introduces discrimination or regulatory violations. Fannie Mae and Freddie Mac formalized these concerns in 2026 with new AI governance requirements, mandating that lenders manage AI risks and disclose their systems upon request. Without standardized testing, each company defines success differently—potentially creating systemic risk if many rely on similar flawed models. More than a third of U.S. ChatGPT users have discussed personal finances with the chatbot, even as 82% describe those conversations as sensitive, yet consumers have little visibility into how their data is processed.
What borrowers should ask
Diane Yu, co-founder and CEO of Tidalwave, a mortgage technology company that collaborated on the research, advises borrowers to ask lenders whether their information goes to outside language models and whether the AI has undergone independent evaluation. "You should ask those questions," she said. "You should be very careful."
Because MortarBench is open source, it allows consistent evaluation across different AI systems rather than leaving vendors to grade their own work. "If we don't measure it, then we don't know about it," Toles said.
These findings were first reported by Realtor.com.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call