The pixel
The pixel
Every real estate and mortgage company building with AI right now is running into the same wall: the model is only as good as the records underneath it. The Warren Group CEO David Lovins made that case directly on a recent episode of ORIL’s Innovation Blueprint Podcast, hosted by Roman Havrylyuk, in a conversation about what it actually takes to turn public-record data into something an AI system, or a person, can trust. 

The short version of David’s argument: AI does not reduce how much data quality matters. It increases it. Public records originate locally, in inconsistent formats, collected by thousands of different towns, counties, and registries with their own conventions. As AI pushes analysis down to the transaction level, the anomalies that used to average out in a spreadsheet start showing up as confident, wrong answers in a model. Roman shared the conversation on LinkedIn under a title that gets right to the point: AI’s data quality challenge. 

Data is easy to buy. Data expertise is harder to find. 

David’s core distinction in the episode is between selling data and understanding it. Plenty of providers will hand over a bulk file. Far fewer can tell you where that data actually originated, how it was collected, what had to be normalized to make it usable, and where the soft spots are. The Warren Group has been doing exactly that work since 1872, which David sums up as being “small and mighty,” customized enough to work with a single team’s specific use case, experienced enough to know the data inside and out after 150-plus years of handling it. 

That distinction matters more, not less, once AI enters the picture. Companies building AI engines and analytics products are not usually looking for one dataset. They come to TWG needing property, mortgage, and ownership records stitched together in a way that actually holds up under automated analysis. That gives The Warren Group a credible position in the AI conversation without needing to rebrand itself as an AI company. The pitch is simpler than that: AI can build the engine, but it still needs real fuel. 

Transparency is the differentiator, not perfection 

One of the more direct points David made is one a lot of vendors avoid saying out loud: no public-record dataset is perfect. Every source has gaps, quirks, and known limitations, and a responsible data partner does not pretend otherwise. Instead of just claiming to have the best data, David argues the job is to help customers understand exactly where the limitations sit, so they can account for them in whatever they are building on top of it. That transparency becomes especially important as more organizations lean on AI to catch patterns a person might miss. A model trained on data with unexplained blind spots will confidently reproduce those blind spots at scale. 

This is also where a startup-friendly posture pays off. David described a steady influx of bootstrapped companies coming to TWG for multiple datasets at once, and rather than leading with a full-catalog price tag, the team works to understand what the product actually needs first. That might mean starting with a single geography, a prototype, or a proof of concept, and scaling the data relationship as the product grows. It is a solutions-first approach: tell TWG what you are trying to build, not just which dataset you think you need, and the team will help connect the mission to the right data and delivery method, including pointing to a partner’s data when that is the better fit. 

Where the datasets are headed next 

The episode closes on where David sees this heading: from individual datasets toward genuinely connected property intelligence. Right now, a customer working across deed and mortgage recordsassessor data, and permit history has to do real work to tie those sources together into one coherent picture of a property. David’s view is that a common identifier across TWG’s datasets would make that ingestion dramatically easier, both for the humans doing the analysis and for the AI systems trying to reason across records that were never designed to talk to each other. It’s a natural next step for a company whose entire value proposition is making messy public information usable, applied now to making its own datasets easier to connect to one another. 

How The Warren Group Can Help 

If your team is building an AI product, a scoring model, or an analytics layer that depends on real estate or mortgage data, the conversation David described on the podcast is the one worth having before you sign a data contract. That means asking where the data actually comes from, how it gets normalized, and what its known limitations are, not just how many records are in the file. The Warren Group works from property, deed and mortgage, NMLS, HOA, and permit data, and the team’s approach starts with understanding what you are trying to build rather than quoting a flat rate off a data catalog. For startups and teams testing an early product, that includes scoping a proof of concept or a single geography rather than requiring a full national commitment on day one. 

Conclusion 

AI has made a lot of companies rethink what “good data” actually means, and David Lovins’ conversation on the Innovation Blueprint Podcast is a useful, unvarnished look at what that shift looks like from the vendor side of the table. The takeaway isn’t that AI makes data quality less of a concern. It’s the opposite. If you want to hear the full conversation, watch the episode here, and if you’re working through what AI-ready real estate or mortgage data should actually look like for your product, contact our team to start the conversation.