Running ML models in the browser with ONNX Runtime Web: a practical guide

Presenting to a room
Presenting to a room

Originally published at https://pranjulrathour.scult.in/blog/onnx-runtime-web-browser-ml-guide. That copy is the canonical version and gets updates first.

FaceVision does face detection, recognition and liveness checks without sending a single frame to a server, because every model runs in the browser through ONNX Runtime Web. That architecture is the reason its privacy story is simple. It is also the part students find hardest to get working, so here is the path that worked.

Export, then verify

Export the model to ONNX from its training framework and immediately run the same input through both the original and the exported graph. Compare outputs numerically. Every model in FaceVision was verified against its actual ONNX graph and, where possible, the reference implementation — not assumed from documentation. Half the "the model is broken in the browser" reports I have seen were pre-processing mismatches that a five-minute comparison would have caught.

Requirements gathering and user flows, on stage
Requirements gathering and user flows, on stage

Choose the execution provider deliberately

  • WebGPU — fastest where available; check support and fall back gracefully.
  • WASM with SIMD and threads — the reliable default; needs the right headers for multi-threading.
  • WebGL — legacy; avoid for new work.

Pre- and post-processing is where bugs live

Models expect a specific channel order, normalisation and input size. Browser image data arrives as RGBA bytes in a different layout. Write the conversion once, test it against a known image, and keep the output tensors' shapes in a comment next to the code. Post-processing — decoding anchors for a detector, normalising an embedding before cosine similarity — deserves the same care.

Performance that feels live

  1. Warm up the session with one dummy inference on page load; the first run is always slow.
  2. Run inference in a Web Worker so the camera preview never stutters.
  3. Downscale frames before detection; run recognition only on detected crops.
  4. Skip frames when the queue is full rather than letting latency grow.

What the backend does when the model does not

In FaceVision the FastAPI and PostgreSQL backend stores only vector embeddings and matches them at enrolment and verification. Because no images are ever persisted, the privacy claim is easy to explain and audit. I wrote about the design reasoning in face recognition that never uploads a face.

Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language

Browser-side inference is not a gimmick. For anything involving faces, documents or health data, it is often the architecture that makes the product acceptable to the people using it.

From my carousels
5 Production AI Apps, All Open Source
5 Production AI Apps, All Open Source, slide 15 Production AI Apps, All Open Source, slide 2
5 Production AI Apps, All Open Source, slide 35 Production AI Apps, All Open Source, slide 4
Full carousel on Instagram and LinkedIn.
Pranjul Rathour
Pranjul Rathour
GenAI engineer, Kanpur · 3x first-prize hackathon winner · campus mentor
I ship production RAG pipelines, fine-tune LLMs and build agentic AI products end to end. I lead engineering at SCULT INDIA for a 14-member team and have mentored 200+ students through TechVerse Enclave.
Open to: GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.
On stage, at hackathons and on campus
Taking questions during a session
Taking questions during a session
Pranjul Rathour
Pranjul Rathour
Presenting Annapurna on stage
Presenting Annapurna on stage

Comments

Popular posts from this blog

Forming a hackathon team: roles, skills and the mistake most teams make

Hello from Kanpur: what I build, and what I'll write about here

I built 15 free tools, 1,211 prompts and a 50,000-skill library — here's what's inside tools.scult.in