Mumbai, Maharashtra, India
Junior Data Scientist shipping production-grade ML in logistics — from delivery date prediction to dispatch intelligence to customer communication. I joined ShipDelight not knowing I'd be helping build the data science function from the ground up. Turns out, the delivery predictions customers rely on every day were outsourced to a third party. We brought them in-house, rebuilt them from scratch, and made them significantly better. That experience taught me more about owning ML systems end-to-end than any course or internship ever did. Today I work across the full spectrum of logistics intelligence. On the prediction side, I've built and improved XGBoost-based EDD systems across multiple timeframes, introduced custom loss functions to improve robustness, and co-developed a LightGBM buffer model that pushed accuracy to 90%+ within ±1 day. I also built an intelligent model-selection system that dynamically routes each shipment to the best-performing model based on historical data. On the analytics side, I built DispatchIQ — a 6-module dispatch intelligence system that gives clients visibility into dispatch delays, GMV at risk, live shipment risk, SKU-level friction, and location-based delay patterns. And on the communication side, I'm developing systems that decide when to reach customers and how to speak to them in a way that actually feels human. All of this runs on 20–30K shipments every day. Outside of logistics, I've fine-tuned LLaMA2 for multilingual image prompt generation across Indian languages, achieving 60%+ improvement in prompt-to-image quality. I also built Therapy Connect — a full-stack AI application with an end-to-end pipeline combining Whisper-based audio transcription, a custom BiLSTM emotion detection model, and a RAG system using ChromaDB and Gemini Pro to generate contextually grounded therapeutic responses. I'm particularly drawn to problems where ML directly changes what a real person experiences — whether that's knowing when their order will arrive or feeling heard in a difficult moment. Current Stack: XGBoost · LightGBM · LSTM · Airflow · Python · MongoDB · AWS (S3 · EC2 · CloudWatch)
• As part of ShipDelight's first in-house data science team, co-led migration of EDD prediction from 3rd-party vendor to in-house ML. • Introduced Pseudo-Huber Loss and asymmetric loss functions into XGBoost EDD system (3 timeframes: 3M, 1M, 15D) — achieving 75%+ accuracy within ±1 day. • Co-developed LightGBM buffer model (regressor + classifier) on XGBoost outputs, pushing accuracy to 90%+ within ±1 day. • Built intelligent model-selection system that automatically identifies and routes each shipment to the most accurate XGBoost timeframe model (3M, 1M, 15D) using historical performance data. • Built DispatchIQ — a 6-module dispatch intelligence system that surfaces dispatch window distribution, DRI scoring, and live GMV at risk for logistics clients. • Developed real-time shipment risk tagging (Low/Medium/High), SKU-level friction scoring, and city/pincode delay intelligence. • Designed Airflow DAGs to automate retraining pipelines for XGBoost EDD and DispatchIQ systems — including database updates and automated accuracy reporting via email. • Developing LSTM-based EDD recalibration system to dynamically update delivery estimates when operational delays occur mid-shipment. • Designed end-to-end LSTM architecture and feature pipeline; built regressor to predict remaining ETA and co-developed classifier to estimate delivery probability across same-day, next-day, and delayed outcomes. • Developing Communication Controller that intelligently decides when and what to communicate to end-customers — with built-in spam suppression to ensure only high-priority delivery updates reach the customer. • Developing Humanised Tracking system that translates raw shipment events into friendly, jargon-free customer messages — replacing operational terminology with personalised, context-aware delivery updates.
• Worked with technology and data teams on analytical workflows, production systems, and internal operational tools. • Gained exposure to large-scale logistics workflows, databases, and ML-driven operational systems.
• Built Retrieval-Augmented Generation (RAG) pipelines using OpenAI models and alternative LLM approaches. • Worked with LlamaIndex for document retrieval, indexing, and AI workflow integration. • Performed data preprocessing, retrieval optimization, and experimentation for AI-based applications. • Assisted in improving existing AI systems through testing, evaluation, and iterative enhancements.
My internship at Final Project provided invaluable practical experience across diverse machine learning domains. I honed my skills, tackled complex challenges, and made impactful contributions to data-driven decision-making, optimizing machine learning models.
My internship at ineuron, a data science leader, enriched my skills. It built a strong data science foundation, empowering me in complex problem-solving, machine learning, and impactful data-driven decision-making.