Document intelligence for lending paperwork
One pipeline that learns a new document type, classifies what arrives, extracts the fields operations needs, and says how confident it is about each one.
- Sector
- Housing finance, operations and legal
- Period
- 2025 to 2026
- The number
- Per-field confidence on every extraction
The leak
A home loan is paperwork. Cheques, printed and handwritten. Demand drafts and bank deposit slips. Salary slips and account statements. Sale deeds, powers of attorney and lease agreements, often in a regional language. Passbooks, rent receipts, utility bills. Every one of them is read by a person who types a few fields from it into a system, and the legal team alone has hundreds of document types that matter to a title or a disbursement.
The reading is slow, the typing is error-prone, and the same document is often read twice by two teams who need different fields from it. None of that reading leaves anything behind: the next time a similar document arrives, the work starts from zero.
The constraint
The documents are varied in a way that defeats a single model. Handwriting on a cheque and a typed sale deed in a regional script are different problems. Many legal documents are scanned at low quality and run to dozens of pages. And the output has to be trustworthy enough for operations to act on, which means the system has to say not just what it read but how sure it is, field by field, so a person checks the uncertain ones and trusts the rest.
It also had to be a pipeline, not a project. Adding the next document type could not mean building the next system.
The system
A single pipeline with a defined way to onboard a new document type: collect examples, define the fields that matter and what they look like, train the classifier to recognise the type, and configure extraction for its fields. Every document that flows through is classified, routed to the extraction configured for its type, and returned as structured fields with a confidence score on each. All examples and corrections collect in one place, so the models improve from the work operations is already doing.
For the hardest class, property legal documents, the pipeline renders each page at high resolution, cleans it, and converts it to structured text before extraction, handling Hindi, Marathi, Gujarati and English. From a sale deed it returns the parties, the registration details and office, the property's address, area and boundaries, the consideration, and a confidence per field, in English, with an audit log of what it read and why.
The models underneath are open vision and language models, repurposed and tuned for these documents and run on the company's infrastructure, so the data never leaves and the cost is compute rather than a vendor's meter. The output feeds the systems that need it: the title research system reads the deed fields it extracts; the underwriting assistant will read statements and slips through it.
The number
Every extraction carries a per-field confidence, so operations reviews what the system is unsure about and trusts what it is not. The first document types are in production, cheques, slips, statements and deeds, with the onboarding process now the way the next hundred are added.
What I would do differently
Build the correction loop into the operations screen on day one. Corrections are the training data, and when they are captured where people already work, the pipeline improves without anyone being asked to label anything. That loop arrived after the first types were live, and the types that came after it improved faster.