Seoul Incheon Metropolitan Area
* Developing real-world video understanding model
∗ Research and developed large vision language model, especially focusing on improving reasoning ability by reinforcement learning ∗ Maintained training LVLM codebase used by 10+ collaborators ∗ Built multilingual CLIP models of varying scales and loss architectures leveraging 100+ GPUs for distributed training
∗ Developed training pipeline for post-OCR and OCR-free parser ∗ Conducted research on efficient information extraction methods from pretrained backbones
∗ Analyzed text and quantitative data for Fraud Detection System(FDS) ∗ Developed post-OCR parser ∗ Built customer-facing question answering system using embedding-based retrieval model