Kartik Vij

Property title research, automated

The legal team's thirteen-year ownership search across state registry portals, run by a system that knows every portal step by step.

Sector
Housing finance, legal
Period
2025 to 2026
The number
Hours per case to minutes

The leak

Before a home loan is disbursed, the lender's legal team has to be certain that the property belongs to the person selling it, that nobody else has a claim on it, and that it has not been quietly sold or mortgaged along the way. In practice that means pulling every registered transaction on the property for the last thirteen years from the relevant state's registry portal, reading each deed, and assembling the chain of ownership into a title report that a lawyer signs.

For a lender working across several states, every state is a different portal with its own search forms, its own way of paginating results, its own document formats and its own failure modes. A lawyer would spend hours on a single case, and the honest accounting of those hours was uncomfortable: most of them went to navigating forms, waiting for pages that stalled, retyping CAPTCHAs, downloading PDFs one at a time and copying details into the file. Very little of it was legal judgement.

Multiply that by the case volume of a growing lender and the leak is a team's worth of time every month. It also shows up on the customer's side as days added to every disbursement, because the title check sits on the critical path.

The constraint

Government registry portals are built for a person with a browser, not for a system. None of them offer an API. Each has its own CAPTCHA format, and several require a one-time password sent to a registered phone before a search can run. They rate-limit and sometimes block an address that queries too often. They stall under load in ways that look like success until you read the page. And they change without notice.

The second constraint was trust. The output of this system is a document a lawyer signs their name to. It could not be a black box. Every record it found, every comparison it made and every place it was unsure had to be visible, so the lawyer reviews evidence rather than taking a system's word.

The third was the rule I set for myself: the model does the parts that need judgement and nothing else. Navigation is deterministic, because a portal is a defined process and improvising on a government website is how you get blocked.

The system

A browser automation that knows each state's portal as a scripted sequence: how to reach the search form, how to fill it from the loan file, how to page through results, how to open and save each document, and how to tell a stalled page from a slow one. Each state is its own module, written once by reading the portal the way a lawyer would, and versioned so that a portal change is a change to one module.

Two steps in that sequence needed something more than a script. CAPTCHAs differ by portal and most vision models fail on them as presented, so each image is pre-processed first: cleaned, thresholded, cropped and normalised until a small, inexpensive vision model reads it reliably. The pre-processing is what made the model cheap enough and accurate enough to run on every case. One-time passwords are routed through the system to a person on the legal team, who enters the code in the same screen; the automation never holds a phone it should not have.

Once the records are retrieved, a language model compares each one against the loan file: party names with their spelling variants, addresses, survey and plot identifiers, dates, consideration amounts. Matches are confirmed with the evidence beside them. Mismatches are explained rather than hidden. Any transaction in which the current applicant appears as a seller in an earlier deed is flagged as an adverse entry for a lawyer to look at first.

Around that core: a job queue with retries so a stalled portal is a delay rather than a failure, archived screenshots and documents for every run, role-based access and the company's single sign-on, and an audit trail that shows who ran what and when. The team opens a case, points the system at it, and reviews the result.

The number

Hours per case down to minutes, and for a clean property with few transactions, seconds. The legal team uses the system every working day. Lawyers' time now goes to the records that need a lawyer: the ambiguous names, the adverse entries, the chains with gaps.

What I would do differently

Two things. The first ceiling we hit was rate limiting: a single machine making many requests to a portal gets throttled or blocked. The design I am moving to is a pool of small workers, including idle office computers running a minimal Linux image, that the orchestrator distributes jobs across. I would build that pool first, because it changes how the whole system schedules work.

The second is integration. Today the team runs the system from its own screen and brings the result back into the loan origination system. The next version runs inside it, so the title check starts when the file is ready and the lawyer never leaves the screen they work in. I would have built that bridge from the start.