Kartik Vij

Why we moved KYC checks on-device

Blur, OCR, masking, liveness and face match were all paid API calls. Most of them run fine on a phone. The one that does not runs on our own server.

Customer onboarding at a lender runs on identity documents. A customer photographs their ID, takes a selfie, and a chain of checks runs before anything is stored: is the photo usable, what does it say, which parts must be hidden, is the person real, is it the same person.

When I arrived, every link in that chain was a vendor API. A call for blur detection. A call for text recognition. A call for masking. A call for liveness. A call for face match. Each one billed per use, so the onboarding cost grew with every customer, which is the wrong direction for a cost to grow in.

The question was not whether these checks could be done without vendors. It was which ones could be done on the device the customer is already holding.

Most of them, it turns out.

Blur detection is a variance calculation on an image. It runs on the phone in milliseconds and stops a bad photo before it is uploaded, which also saves the round trip.

Text recognition on identity documents is a solved problem on-device. The standard mobile toolkits read the fields well, and a small model trained to recognise the handful of entity types we care about picks out the right ones. Masking follows from recognition: once you know where the digits are, covering the ones you must not keep is drawing a rectangle. Watermarking is the same operation in reverse.

Liveness was the surprising one. The vendor approach is passive: analyse the selfie for signs of a screen or a print. We replaced it with an active check. The app asks for something random, turn your head left, now blink, now look right, and captures the frame at the moment the person has just complied. It is harder to fake, it needs no model, and it gives the next step a better photo.

Background removal for that selfie, so the stored photo is on plain white, runs on-device too.

Then face match, which is the real model. Comparing a selfie to the photo on an identity document, across lighting, age and camera quality, is not something a low-end phone does well. So it stayed on a server. But not a vendor's server. Open face-recognition models, repurposed and tuned on the kinds of documents we actually see, run on machines the company already pays for. The cost of a match became the cost of the server, which does not change when the hundredth customer becomes the thousandth.

A few things we learned.

The low end of the market sets the floor. Everything had to run on the cheapest Android phone the customer base uses, which ruled out anything heavy. That constraint made the design better, not worse.

Masking has rules, not preferences. Which digits of a national identity number may be stored is defined, so the masking had to be exact. Approximate was not an option.

Measure against the vendor from the first week. The one thing I would change is that the accuracy comparison came later than it should have. Each component should carry its own proof on the day it is proposed, on the same test set as the service it is replacing, so switching the vendor off is a decision with evidence rather than a leap.

The chain is being switched off one link at a time, as each on-device component earns it. The per-call bill is heading to zero for everything except the one step that genuinely needs a model, and that step costs a server.

If a check is a calculation, it belongs on the device. If it needs a model, it belongs on your server. It only belongs on a vendor's meter when neither of those is possible.