San Francisco, California, United States
portfolio: https://mahfooz.me/
compute architecture
running lab sessions for EECS1022 (intro to OOP) and EECS1520 (computer use: fundamentals)
compute architecture • Developed a suite of tools to deploy hardware-agnostic SLURM 22 clusters, enabling teams to test and iterate on emerging GPU/CPU/driver/OS configurations more efficiently. • Designed and built a GPU tracing pipeline on Kubernetes and built supporting infrastructure to capture and analyze GPU/CPU/system performance metrics for deep learning inference. • Developed microservices and infrastructure that interface with SLURM and other internal tools to classify DL network topologies, collect performance metrics, and streamline cluster management.
enterprise architecture & technology enablement • Developed a reference implementation for rapid microservice deployment on AWS EKS, utilizing Spring and Kafka, aiding in the transition from legacy monolithic services by reducing repetitive setup work for developers. • Utilized OAuth 2.0 for securing API calls, migrating from Apigee APIM, improving user security for 13,000,000+ clients. • Implemented automated GitLab CI/CD pipelines for building, containerizing, and deploying microservices on EKS.
middleware • Designed and implemented an end-to-end transcription & insights service for incoming customer service calls using OpenAI Whisper, hosted on an Azure Databricks GPU cluster, resulting in enhanced call analysis accuracy and efficiency. • Employed advanced NLP models/libraries including GPT-4, BERTopic, and spaCy to develop a microservice that extracted key insights and information from customer calls, enabling data-driven decision-making and improved customer support. • Established a robust microservice architecture to orchestrate communication between the GPU-clustered transcription engine and the AI-driven insights extraction, ensuring high scalability and reliability for the transcription & insights service. • Successfully migrated internal chatbot from using Azure Cognitive Services to Azure OpenAI GPT-4, improving response times by up to 200% and response quality by up to 300% for over 1300 employees. • Developed backend HTTP/1.1 endpoints with Flask, Redis (vector DB), and data stores (Snowflake, Azure Blob Storage). • Leveraged Azure Kubernetes Service (AKS) for scalable container orchestration and implemented a streamlined CI/CD pipeline with GitHub Actions, Jenkins, and K8s for seamless deployment and management of the chatbot application.