Designing an OCR pipeline: from scanned PDF to searchable, chunked text

Requirements gathering and user flows, on stage
Requirements gathering and user flows, on stage

Originally published at https://pranjulrathour.scult.in/blog/ocr-pipeline-design-pdf-to-searchable-text. That copy is the canonical version and gets updates first.

The OCR & Speech Workspace was built to make books searchable — page-batched, concurrent OCR over PDFs, then document-scoped RAG chat with page citations. The OCR model is the easy part. The pipeline around it is what decides whether page 214 is findable six months later.

Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language

Split into pages, then batch

Process the PDF page by page and OCR pages in concurrent batches. Per-page processing gives you natural checkpoints (a crash on page 300 does not lose 299 pages), natural citations, and natural parallelism. Tune batch size to the OCR provider's rate limits, and record per-page status so a re-run only touches failures.

Keep the layout, not just the words

Modern OCR returns structure — headings, paragraphs, tables — as markdown or blocks. Preserve it. A table flattened into a paragraph is unsearchable and uncitable. Structure is also what makes chunking for RAG work downstream: heading context travels with each chunk.

Taking questions during a session
Taking questions during a session

Quality checks that catch bad pages

  • Character count far below the document average — a blank or failed page.
  • High ratio of non-dictionary tokens — a rotated or low-resolution scan; re-run with rotation detection.
  • Repeated headers and footers — strip them before chunking or every chunk starts with the book title.
  • Language detection per page for bilingual documents, so the right embedding model is used.

Store for search and for citation

  1. Text per page, with page number and document id.
  2. Chunks with page ranges and heading context, embedded for hybrid retrieval.
  3. A content hash per document so re-uploads do not re-OCR.

What it enables

Once pages are text with metadata, "where does this book discuss X?" becomes a retrieval question with a page-level answer, and citations in RAG answers become links a reader can open. I wrote about the workspace itself in every page and every word, searchable.

From my carousels
5 Production AI Apps, All Open Source
5 Production AI Apps, All Open Source, slide 15 Production AI Apps, All Open Source, slide 2
5 Production AI Apps, All Open Source, slide 35 Production AI Apps, All Open Source, slide 4
Full carousel on Instagram and LinkedIn.
Pranjul Rathour
Pranjul Rathour
GenAI engineer, Kanpur · 3x first-prize hackathon winner · campus mentor
I ship production RAG pipelines, fine-tune LLMs and build agentic AI products end to end. I lead engineering at SCULT INDIA for a 14-member team and have mentored 200+ students through TechVerse Enclave.
Open to: GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.
On stage, at hackathons and on campus
Pranjul Rathour
Pranjul Rathour
Presenting Annapurna on stage
Presenting Annapurna on stage
Presenting to a room
Presenting to a room

Comments

Popular posts from this blog

Forming a hackathon team: roles, skills and the mistake most teams make

Hello from Kanpur: what I build, and what I'll write about here

I built 15 free tools, 1,211 prompts and a 50,000-skill library — here's what's inside tools.scult.in