Gyeonggi, South Korea
Machine Learning Engineer with 7+ years of experience designing, training, and deploying large-scale ML systems for search and recommendation. I've developed state-of-the-art multimodal models trained on 1000+ GPUs over 800M+ items, and built the audio fingerprinting technology behind NAVER's major services (N-Innovation Award, 2019). My work spans the full ML lifecycle — from research and modeling to distributed training infrastructure and production deployment. Passionate about driving innovation and delivering impactful AI solutions at scale.
Focused on developing and deploying large-scale machine learning solutions for NAVER's Shopping Search. - Led the development of SOTA multimodal foundation models (CLIP variants) tailored for shopping search, achieving 10+% mAP and nDCG improvements through techniques like momentum distillation and multi-task learning. - Orchestrated large-scale distributed training on 1024 A100 GPUs for ~800M image-text pairs, optimizing data pipelines (HDFS/DDN) to achieve 98% GPU utilization. - Created the "omni-lightning" experimentation framework (Hydra) to accelerate research and development cycles. - Built an internal UI (Nuxt.js) for multimodal search analysis. - Engineered and deployed a high-performance recommendation API (CLIP embeddings + SIFT) handling 8000 QPS with <100ms latency - Implemented robust MLOps pipelines (Kubeflow, Airflow) for automated daily embedding generation (~40M items/day), monthly full generation (1.6B items / monthly), ANN indexing (Faiss), and model updates. - Optimized model serving for production using TensorRT/Triton, achieving <10ms inference latency on A100 GPUs. - Invented a patented method for multimodal feature manipulation (feature arithmetic) to enhance search capabilities. - Developed a high-accuracy (98%) ML pipeline for fashion color extraction using YOLO, U^2Net, and k-means.
Focused on developing core audio analysis and search technologies for NAVER's music services. - Invented and developed a novel, noise-robust audio fingerprinting algorithm, achieving 90% recall @ 95% precision on a 10M song database. This technology currently powers NAVER & LINE music search. - Received the NAVER N-Innovation Award (2019) for the audio fingerprinting technology, recognizing top innovations within the company. Secured multiple international patents for this work. - Developed a cover song detection system using metric learning and CNNs, achieving 85% accuracy on a challenging non-melodic dataset. - Mentored two KAIST interns on research projects related to Query-by-Humming (DTW optimization) and audio classification, leading to a patent filing.