Mohit Deharkar

CV Intern @Vehant | Pre-final year @IITJ

Nagpur, Maharashtra, India

About

Experience

  • Computer Vision Intern at Vehant Technologies
    May 2026 - Jul 2026 · 3 mos

    Worked under Dr. Shikha Gupta at the R&D Labs. • Implemented and evaluated classical motion segmentation approaches (MOG2, ViBe, and Hybrid models) for real-time video streams, analyzing performance limitations such as ghosting and illumination sensitivity. • Improved the ViBe background subtraction algorithm using adaptive update rates and temporal quantization strategies, reducing background artifacts while maintaining 30 FPS edge inference with 3× lower memory usage ( 10 MB). • Developed a real-time waterlogging segmentation pipeline for Western Coalfields Ltd. (WCL), benchmarking ResNet- UNet ( 106 MB) against SegFormer architectures ( 51 MB) and optimizing FP16 deployment to achieve 6× model com- pression, 48% lower GPU memory usage, and 1.43× faster inference with < 0.14% prediction variation.

  • GeoSpatial NLI — Multimodal Vision-Language System for Satellite Imagery at Inter IIT Tech Meet 14.0
    Nov 2025 - Dec 2025 · 2 mos

    Built an end-to-end vision–language pipeline for an ISRO problem statement that allows users to query satellite images in natural language and receive captions, answers, and grounded object locations. 🔹Designed a multi-modal pipeline for RGB, IR, SAR, and FCC imagery with automatic modality classification 🔹Implemented captioning, VQA, and oriented object grounding using open-source VLMs (Qwen-VL, Moondream2) 🔹Solved SAR language grounding without SAR captions by integrating SARATR-X + LLM-based spatial reasoning 🔹Enabled scale-robust inference on high-resolution imagery (up to 2k×2k, 0.5–10 m/px) 🔹Deployed the complete system as a full-stack web platform (Django backend + React frontend) 🔹Focus Areas: Vision–Language Models · Remote Sensing · SAR · Multimodal AI · Spatial Reasoning