Judging AI hackathon projects: what to check when every team says 'we used AI'

Trophy and certificate of merit at Vividhotsava 2025
Trophy and certificate of merit at Vividhotsava 2025

Originally published at https://pranjulrathour.scult.in/blog/judging-ai-hackathon-projects-what-to-check. That copy is the canonical version and gets updates first.

Half the projects at any student hackathon now include an AI component, and most panels have no reliable way to tell a real one from a wrapper around a prompt. Having built RAG, fine-tuning and vision systems for a living, these are the checks I run in three minutes at a judging table — and the ones I share with fellow judges beforehand.

Is the inference real?

  • Ask for an input the team did not prepare. A hard-coded response will not adapt.
  • Ask to see the request go out — a network tab, a log line, a token counter. Real calls leave traces.
  • Ask what the model was given as context. Teams who built it can answer instantly.

Did they evaluate anything?

"We tested it on a few examples" is acceptable at a hackathon if they can show the examples and the outcomes. A team with ten test questions and a note on which three failed is showing engineering maturity; score it. A team that claims 95% accuracy with no test set is showing a slide.

Walking a room through evaluation criteria
Walking a room through evaluation criteria

How does it fail?

Ask what the system does when the model is wrong, slow or unavailable. A confidence threshold, a fallback, a clear error message or a human hand-off all count. "The AI handles it" does not. The human decisions around the model are what distinguish a product from a demo.

Where does the data go?

Any project touching faces, health, documents or children should be able to say what leaves the device, what is stored and for how long. A team that thought about this deserves credit even if the implementation is thin.

Presenting BrandHive
Presenting BrandHive

Is the AI the right tool?

Some of the best projects use AI for one narrow step and plain code for the rest. Some of the weakest bolt a chatbot onto a problem that needed a form. Score appropriateness, not quantity of AI.

Scoring it fairly

  • Do not reward the team with the biggest model; reward the team with the clearest understanding.
  • Do not penalise a scoped-down but real inference against a broad but mocked one.
  • Write the AI-specific feedback down; it is the part teams most need and least often get.

Organisers running AI-themed hackathons: brief your judges with this list, or invite one who already uses it. I am open to judging student hackathons on-site or remotely; details on the campus invite page.

From my carousels
Three First Prizes, One Loss
Three First Prizes, One Loss, slide 1Three First Prizes, One Loss, slide 2
Three First Prizes, One Loss, slide 3Three First Prizes, One Loss, slide 4
Full carousel on Instagram and LinkedIn.
Pranjul Rathour
Pranjul Rathour
GenAI engineer, Kanpur · 3x first-prize hackathon winner · campus mentor
I ship production RAG pipelines, fine-tune LLMs and build agentic AI products end to end. I lead engineering at SCULT INDIA for a 14-member team and have mentored 200+ students through TechVerse Enclave.
Open to: GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.
On stage, at hackathons and on campus
Presenting KrishGyan, farming advice in your voice and language
Presenting KrishGyan, farming advice in your voice and language
Speaking at a MeetKats event
Speaking at a MeetKats event
At an Integral Startup Foundation hackathon
At an Integral Startup Foundation hackathon

Comments

Popular posts from this blog

Forming a hackathon team: roles, skills and the mistake most teams make

Hello from Kanpur: what I build, and what I'll write about here

I built 15 free tools, 1,211 prompts and a 50,000-skill library — here's what's inside tools.scult.in