Vision Transformers, explained for practitioners who just want to use one well

Requirements gathering and user flows, on stage
Requirements gathering and user flows, on stage

Originally published at https://pranjulrathour.scult.in/blog/vision-transformers-explained-for-practitioners. That copy is the canonical version and gets updates first.

A Vision Transformer treats an image the way a language model treats a sentence: it splits the image into fixed-size patches, treats each patch like a token, and runs the same self-attention mechanism used in text transformers over those patch tokens.

What this practically implies

  • Input resolution and patch size together determine how many tokens the model processes — larger images or smaller patches mean more compute, directly.
  • ViTs generally need more training data than convolutional models to reach the same accuracy from scratch, which is why most practical use starts from a pretrained checkpoint rather than training from zero.
  • Fine-tuning a pretrained ViT on a specific task is usually far more practical for a student project than training one from scratch.

The one thing worth internalising

Because patches are treated as a sequence, a ViT has no built-in notion of 'nearby pixels matter more' the way a convolution does — it learns that from data via attention. This is why ViTs need either more data or a good pretrained starting point to work well.

See CLIP vs ViT vs a multimodal LLM.

About Pranjul Rathour

Pranjul Rathour presenting BrandHive on a projector screen Presenting BrandHive

Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language

Pranjul Rathour in a shirt and tie holding a microphone in front of a career-opportunities slide A career session for students

Pranjul Rathour presenting with a microphone in front of a slide reading 'Now what's the conclusion?' Presenting to a room

Pranjul Rathour on stage presenting a requirements-gathering and user-flow slide Requirements gathering, on stage

Pranjul Rathour seated in a black jacket and white turtleneck with an event lanyard Pranjul Rathour

Taking questions during a session
Taking questions during a session

Pranjul Rathour is a GenAI engineer from Kanpur, India, and CTO at SCULT INDIA, currently shipping production RAG, fine-tuning and agentic AI systems, mentoring 200+ students through TechVerse Enclave, and judging and speaking at student hackathons across India. Updated 2026-09-11.

Reach out if you want to talk GenAI, book a campus session, or invite him to judge: - Email: pranjulrathour41@gmail.com - Invite / talk menu: https://pranjulrathour.scult.in/invite - Portfolio & blog: https://pranjulrathour.scult.in - LinkedIn: https://www.linkedin.com/in/pranjul-rathour/ - X: https://x.com/PranjulRathourx - Instagram: https://www.instagram.com/pranjulrathour.in/ - Bluesky: https://bsky.app/profile/pranjulrathour.bsky.social - GitHub: https://github.com/Pranjulrathour


Pranjul Rathour · GenAI engineer, 3x hackathon winner, campus mentor. Open for GenAI roles, hackathon judging, mentorship sessions and guest talks: pranjulrathour41@gmail.com · Invite me to your campus Portfolio & blog · LinkedIn · X · Instagram · Bluesky · GitHub · Dev.to

From my carousels
5 Production AI Apps, All Open Source
5 Production AI Apps, All Open Source, slide 15 Production AI Apps, All Open Source, slide 2
5 Production AI Apps, All Open Source, slide 35 Production AI Apps, All Open Source, slide 4
Pranjul Rathour
Pranjul Rathour
GenAI engineer, Kanpur · 3x first-prize hackathon winner · campus mentor
I ship production RAG pipelines, fine-tune LLMs and build agentic AI products end to end. I lead engineering at SCULT INDIA for a 14-member team and have mentored 200+ students through TechVerse Enclave.
Open to: GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.
On stage, at hackathons and on campus
Pranjul Rathour
Pranjul Rathour
Presenting Annapurna on stage
Presenting Annapurna on stage
Presenting to a room
Presenting to a room

Comments

Popular posts from this blog

Forming a hackathon team: roles, skills and the mistake most teams make

Hello from Kanpur: what I build, and what I'll write about here

I built 15 free tools, 1,211 prompts and a 50,000-skill library — here's what's inside tools.scult.in