LLMs LLM Post-Training on Small-Scale Models SFT, RLHF with process reward models, and latent-space knowledge distillation for improving mathematical reasoning in Qwen2.5-0.5B. Multi-Step Agentic RAG System for Multi-Hop QA Decoupled agentic RAG pipeline with multi-step retrieval over 270K facts from HotpotQA. FastAPI Safety Gateway Guardrails + agentic safety loop for a vLLM-served Llama-3.1-8B-Instruct. Datasets Video Humor Reasoning (Indic-SMILE) Dataset + benchmarking LLMs on culturally grounded Hindi humor understanding. Competitions Deepfake Image Detection Challenge Led data + model evaluation for an open-source deepfake detection effort (Omdena); ranked 3rd.