Santa Monica, California, United States
Innovative ML engineer with 10+ years of experience building and deploying scalable, real-world AI systems. Specializing in conversational AI, speech technologies (ASR/TTS), and LLM-powered solutions, I develop cutting-edge models that enhance human-computer interactions. With a background spanning machine learning, affective science, and music, I thrive at the intersection of technology, psychology, healthcare, and music tech —leveraging computational models to solve complex, multidisciplinary challenges. 🔹 Proven Impact: Delivered AI products used by millions at Cisco and high-growth startups. 🔹 Expertise: End-to-end ML system design, deployment, and optimization in production environments. 🔹 Passion: Building AI-driven speech and conversational systems that push the boundaries of human-AI interaction + music technology for creatives Always open to connecting with like-minded professionals, discussing AI innovations, and exploring new opportunities in AI-driven product development.
Provided technical leadership at the intersection of audio, DSP and machine learning - Built comprehensive training and evaluation framework for music source separation models using PyTorch, Lightning, and Weights & Biases for experiment tracking - Improved stem source separation quality by 4-point increase in Signal to Distortion Ratio compared to previous baseline by implementing BandSplit Roformer model - Developed an audio stem alignment algorithm 100x faster than previous approach, removing critical processing bottleneck - Technologies: PyTorch, Lightning, Wandb, Paperspace, Roformer, Signal Processing
Led AI strategy and team development for art curation platform - Defined and executed product roadmap alongside co-founder to deliver AI-augmented search engine - Led a team of 5 engineers to successfully deliver a cloud-based AI-augmented search engine featuring semantic search, automatic tagging and visual similarity matching using CLIP embeddings - Increased art curator productivity by 3x through intelligent content discovery and organization - Technologies: ElasticSearch, CLIP, Google Cloud Platform, Vector Embeddings
Led end-to-end development of real-time voice agent platform integrating cutting-edge speech and language models for "digital humans" - Architected asynchronous streaming system for real-time voice conversation with end-pointing and interruption handling - Built a speech transcription evaluation pipeline and reduced WER by 35% by using echo cancellation and speech enhancement for various microphone qualities - Implemented a system that leverages user-specific data to support keyword boosting for real-time transcription - Built a speaker embeddings (t-vectors) training pipeline achieving 98% speaker recognition accuracy - Evaluated and optimized voice conversation service latency using OpenTelemetry and Locust to achieve real-time interactions - Technologies: Python, Asyncio, Websockets, Docker, AWS, S3, OpenTelemetry, Locust
Led ML research and development for Webex AI features used by millions of users globally - Gained 4X increase in productivity and number of experiments per week through co-creation of ASR training and evaluation pipeline with Airflow, enabling rapid model iteration and A/B testing - Pioneered Cisco’s first speech synthesis engine integrated into Cisco’s Vidcast video sharing app to provide a novel audio editing feature. Implemented and evaluated various architectures such as Tacotron2, HiFi-GAN, DDPM and Neural Codec LM. - Integrated Voicea.ai captioning into Webex Assistant, providing accessibility features for millions of global users - Developed and deployed 2 hybrid ASR systems for Spanish and French from concept to production, incorporating user-specific vocabulary adaptation - Developed multiple new features and products including the first speaker diarization system (based on x-vector embeddings) and an FST-based on-the-fly decoder with a user-specific vocabulary feature for Webex - Awarded 3 patents for innovations in speech AI and machine learning - Technologies: PyTorch, Kaldi, Kubernetes, AWS, Airflow, FST decoders, Transformers, Diffusion
"From a pool of over competitive 2,800 applicants, we went through a tough decision process of selecting 15 TAs to join the inaugural cohort- 0.5% of the applicants. We designed the TA Program to bring on exceptional new music producers & songwriters seeking to become future leaders of the music production industry"
First ML hire, driving core ASR technology that led to Cisco acquisition - Decreased Word Error Rate by 20% through comprehensive ASR improvements, outperforming competitors and driving acquisition interest - Implemented user and domain adaptation for ASR engines deployed on Kubernetes production clusters using Kaldi and PyTorch - Built automated pipeline for language and acoustic model updates based on labeled data using transformer-based lexicon generation and LLM rescoring as well as on-the-fly addition of OOV - Designed and deployed voice activity detection and speaker diarization services using AWS S3 and SQS, improving meeting transcription accuracy - Technologies: Kaldi, PyTorch, Kubernetes, AWS (S3, SQS)