ANSH SHAH

Software & Machine Learning Engineer | Distributed Systems, Cloud & ML | Python, Spark, AWS, Kubernetes | Building scalable software and data-driven AI solutions

Chicago, Illinois, United States

About

I didn’t start out trying to “build AI at scale.” I started by learning how messy, real-world data can be transformed into reliable software that actually works in production. That curiosity led me into software engineering, distributed systems, and applied machine learning - where scalability, performance, and correctness matter. I’m currently pursuing a Master’s in Computer Science at the University of Illinois Chicago (GPA: 3.8) and have hands-on experience building cloud-native software, data pipelines, and ML systems across academic and industry settings. My work spans software engineering, data engineering, and ML systems, with a strong focus on end-to-end system design. What I work on - Distributed systems and cloud platforms using Python, Scala, Spark, Hadoop, Flink, AWS, Docker, and Kubernetes - ML and NLP pipelines, including retrieval-based systems and analytics workflows - Designing clean APIs, reproducible pipelines, and production-ready infrastructure Selected impact & achievements - Built an end-to-end distributed RAG and GraphRAG platform on AWS using Hadoop, Spark, Delta Lake, Flink, Neo4j, and EKS, processing 600+ research PDFs and exposing 6 REST microservices - Implemented an incremental DeltaRAG indexing pipeline, embedding only new or changed documents and achieving ~30% faster updates while preserving deterministic IDs - Validated large-scale pipelines with 10+ unit tests, structured logging, and runtime metrics including 50,784 token embeddings, ~560 chunks/min throughput, and 85% analogy accuracy - Developed ML models for early Parkinson’s detection, achieving 83% accuracy (Logistic Regression) and 85% accuracy (ANN) using PCA-based dimensionality reduction - Built a full-stack admission prediction platform serving 15,000+ records, achieving 77.9% accuracy across multiple ML models - Designed 5+ Tableau dashboards to support KPI-driven business decisions Experience I’ve worked as a Machine Learning Intern, Data Science Intern, and Software Developer Intern, contributing to model development, analytics pipelines, workflow automation, Kubernetes troubleshooting, Git-based collaboration, and CI/CD improvements. What I’m looking for I’m seeking opportunities to build scalable software and ML systems, collaborate with strong engineering teams, and contribute to data- and AI-driven products with real-world impact. If you’re working on distributed systems, cloud platforms, or intelligent applications, let’s connect. #SoftwareEngineering #MachineLearning #DistributedSystems

Experience

  • Intern at Obmondo
    Jul 2025 - Aug 2025 · 2 mos

    Automated onboarding and offboarding workflows for internal teams at Obmondo, reducing manual effort by 30% and improving operational consistency, using scripted automation and workflow orchestration tools. Resolved 5+ stale Kubernetes alerts across production clusters, improving system reliability and reducing engineering response time, by investigating root causes and applying Kubernetes diagnostic and monitoring practices. Maintained feature branches and resolved merge conflicts using Gitea, strengthening collaboration and release stability during active development cycles, while documenting workflows and improving CI/CD steps for infrastructure readiness.

  • Research Intern at KJ Somaiya College of Engineering, Vidyavihar
    Jun 2023 - Aug 2023 · 3 mos

    Preprocessed a high-dimensional dataset of 756 samples and 755 features at KJSCE, retaining 80% variance through PCA to reduce computational complexity and improve downstream model efficiency using Python-based data processing workflows. Developed supervised ML models for early Parkinson’s detection, achieving 83% accuracy with Logistic Regression and 85% accuracy with an ANN, improving diagnostic reliability through iterative feature evaluation and model tuning. Conducted exploratory data analysis on an imbalanced dataset with 74.6% positive class, identifying skewed distributions to guide model selection and resampling strategies using Python, Matplotlib, and Seaborn visualizations.

  • Machine Learning Intern at SmartKnower
    Sep 2021 - Nov 2021 · 3 mos

    Executed a Student Stress Level Detection project at SmartKnower by preprocessing and analyzing 11,000+ samples with 21 features, enabling accurate classification of stress levels through structured feature engineering and Python-based analysis. Designed and benchmarked multiple ML pipelines using Logistic Regression, SVC, Random Forest, and XGBoost, achieving 87%–94% accuracy and identifying optimal models for robust mental health classification through comparative evaluation. Delivered a production-ready SVC model with 94% accuracy, improving generalizability through class rebalancing and IQR-based outlier treatment, establishing a strong baseline for future extensions in mental health analytics.