How face recognition works: embeddings, thresholds and why you never store faces

Presenting to a room
Presenting to a room

Originally published at https://pranjulrathour.scult.in/blog/face-recognition-embeddings-explained. That copy is the canonical version and gets updates first.

Face recognition is four steps, and only one of them is recognition. Understanding the pipeline is what lets you build one responsibly — which for FaceVision meant a system where the server has never seen a face.

Requirements gathering and user flows, on stage
Requirements gathering and user flows, on stage

The four steps

  1. Detection — find faces in the frame and return bounding boxes and landmarks.
  2. Alignment — rotate and crop each face to a canonical pose using the landmarks, so the next model sees consistent input.
  3. Embedding — a recognition model maps the aligned face to a vector, typically 128 to 512 numbers, such that the same person's faces land close together.
  4. Matching — compare the new embedding to stored ones with cosine similarity and decide with a threshold.

Choosing the threshold

Too strict and real users are rejected; too loose and strangers get in. Collect a small labelled set — pairs of the same person, pairs of different people — and plot the similarity distributions. The threshold sits where the curves separate; where they overlap is your error rate, and you should state it. Lighting, glasses and age change the distribution, so test with your actual users, not a benchmark.

Why you store embeddings, never images

An embedding is enough to verify a person and useless for showing anyone what they look like. Storing only embeddings — the FaceVision design — means a database breach leaks vectors rather than faces, consent conversations get simpler, and deletion is a single row. Run detection and embedding in the browser with ONNX Runtime Web and the raw image never leaves the device at all (see running ML models in the browser).

Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language

Failure modes to design for

  • A printed photo held up to the camera — this is why liveness detection exists.
  • Twins and look-alikes — accept that similarity is not identity and add a second factor for high-stakes actions.
  • Bias — recognition models perform unevenly across skin tones and ages; measure per group before deployment.

Build it this way and you can explain your system to a privacy officer in two sentences. That is a competitive advantage, not just an ethical one.

From my carousels
5 Production AI Apps, All Open Source
5 Production AI Apps, All Open Source, slide 15 Production AI Apps, All Open Source, slide 2
5 Production AI Apps, All Open Source, slide 35 Production AI Apps, All Open Source, slide 4
Full carousel on Instagram and LinkedIn.
Pranjul Rathour
Pranjul Rathour
GenAI engineer, Kanpur · 3x first-prize hackathon winner · campus mentor
I ship production RAG pipelines, fine-tune LLMs and build agentic AI products end to end. I lead engineering at SCULT INDIA for a 14-member team and have mentored 200+ students through TechVerse Enclave.
Open to: GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.
On stage, at hackathons and on campus
Taking questions during a session
Taking questions during a session
Pranjul Rathour
Pranjul Rathour
Presenting Annapurna on stage
Presenting Annapurna on stage

Comments

Popular posts from this blog

Forming a hackathon team: roles, skills and the mistake most teams make

Hello from Kanpur: what I build, and what I'll write about here

I built 15 free tools, 1,211 prompts and a 50,000-skill library — here's what's inside tools.scult.in