Delhi, India
Designed and optimized end-to-end speech-to-text pipelines using Deepgram and Supervox APIs, leading to a 7–8% increase in transcription accuracy across diverse audio inputs. Processed and curated 250+ hours of multilingual and noisy audio data, significantly enhancing model generalization and robustness under real-world conditions. Integrated sentiment analysis modules post-transcription to extract emotional tone from speech, enabling downstream tasks such as customer intent classification and behavioral profiling.