A.Furkan Çomak

Group Platform Manager at Codeway Studios

Barcelona, Catalonia, Spain

About

Experience

  • Codeway (4 yrs 4 mos)
    • Group Platform Manager
      Feb 2026 - Present · 6 mos

    • Senior DevOps Engineer
      Feb 2024 - Feb 2026 · 2 yrs 1 mo

      — Architected and managed high-performance H100 GPU clusters on GKE, specifically optimized for large-scale image generation model training and inference. — Designed and deployed a self-hosted AI Platform using JupyterHub and ClearML to accelerate the ML research-to-production lifecycle. — Founded and led the AI Platform team, serving as the technical and strategic bridge between AI Infrastructure, Development Platforms, and the core DevOps team. — Managed the technical relationship and integration roadmaps with Tier-1 AI vendors, ensuring optimal quota allocation and service reliability. — Managed large-scale cloud governance and FinOps, overseeing multi-million dollar vendor contracts and technical agreements to ensure cost-effective scaling of H100 resources. — Standardized mobile release cycles by building Fastlane automation, reducing manual overhead and increasing the velocity of client builds. — Defined the long-term technical strategy for internal development platforms, prioritizing architectural excellence and eliminating bottlenecks in cross-team engineering productivity. — Pioneered the evolution of the DevOps culture, successfully transitioning the organization from traditional infrastructure maintenance to a highly scalable, platform-centric engineering model. — Optimized GPU resource scheduling for multi-tenant image generation workloads, ensuring maximum hardware utilization and high availability.

    • DevOps Engineer
      Sep 2022 - Feb 2024 · 1 yr 6 mos

      — Designed and implemented CI/CD pipelines for efficient application deployment. — Significantly improved CI/CD pipeline efficiency by reducing processing time by 3x. — Strengthened application security by integrating runtime container security and static code analysis tools into the CI/CD pipeline, ensuring that security measures are an integral part of the development process. — Built and maintained a centralized monitoring system on Kubernetes clusters. — Achieved a significant cost reduction of 80% by transitioning to a self-hosted monitoring platform. — Proficient in managing clusters on various platforms, including GCP, AWS, and on-premises infrastructure, demonstrating strong multi-platform expertise. — Managed over 15 Kubernetes clusters for GPU-based workloads, totaling 500+ GPU-based nodes (mostly A100 and T4 instances). — Reduced cloud compute engine costs by approximately 25% through the implementation of autoscaling strategies. — Optimized GPU utilization and lowered compute spend by implementing advanced resource scheduling and request tuning. — Achieved a 60% reduction in cold-start processing time by significantly improving the ML workload scaling strategy. — Expertise in using PubSub, Firebase, and BigQuery for data processing and real-time communication. — Built and maintained a centralized logging system on Kubernetes clusters. — Managed a centralized alerts and on-call management system. — Proficient in Linux system administration, with a focus on managing on-premises servers, ensuring their reliability and security. — Extensive experience with Nginx server configuration and optimization for robust web server management. — Maintained a centralized user management system within the company. — Led the establishment of a new team within the company, contributing to organizational growth and effectiveness.

  • Data Engineer at Bentego
    Dec 2021 - Apr 2022 · 5 mos

    Consulting Service for Fibabanka Data Transformation Project Responsibilities: — Building & Maintaining Cloudera Distribution Hadoop ecosystem (In place upgrade CDH 6.3 to CDP 7.1.2) — Cloudera Administration — Building Big Data ingestion and processing pipelines (Spark, Spark Streaming, Kafka, Kafka Connect, Sqoop, Airflow, Hive, Impala, MapReduce, YARN) — Building & Maintaining Big Data storage layers (HDFS, Hive, Impala, Kafka, Apache Avro, Apache Parquet) — Building & Maintaining DWH to Datalake/Datalake to DWH Offloading pipelines — Performance Tuning on Spark, MapReduce, and YARN applications — Building & Maintaining Apache Airflow distributed processing (CeleryExecutor) — Implementation & Management of Nexus Repository for artifact & proxy repository — Maintaining & Providing support on JupyterHub Distribution for the DS team — Researching & Development on CI / CD pipelines

  • Data Engineer at ING
    Aug 2021 - Dec 2021 · 5 mos

    — Implementation of Serverless Data Lake Architecture on Openshift/Kubernetes (Spark, Presto, DBT,S3 Object Storage) — New advanced analytics model training environment implementation: JupyterHub on Openshift/ Kubernetes — Implementation of MLOps platform on Azure Devops for models generated on Python batch & web services (PMML & Flask) — Building Big Data ingestion and processing pipelines (Dask, Spark, Spark Streaming, Airflow, Hive)

  • Garanti BBVA Teknoloji (1 yr 4 mos)
    • Big Data System Engineer
      Sep 2020 - Aug 2021 · 1 yr

      — Building & Maintaining Cloudera Distribution Hadoop ecosystem (Oracle BDA to CDH 6.3) — Cloudera Administration — Building Big Data ingestion and processing pipelines (Spark, Spark Streaming, Kafka, Hive, Impala, MapReduce, YARN — Building & Maintaining Big Data storage layers (HDFS, Hive, Impala, Kafka, , Apache Avro, Parquet) — Performance Tuning on Spark, MapReduce, and YARN applications — Maintaining & Providing support on JupyterHub Distribution for the DS team — Building & Maintaining MongoDB and Couch base — Building & Maintaining ELK Stack — Maintaining and project developing in Prometheus and PowerBI

    • Junior System Engineer
      May 2020 - Sep 2020 · 5 mos

  • DevOps Intern at Migros Ticaret A.Ş.
    Oct 2019 - Apr 2020 · 7 mos

    — Creating fully automated CI build and deployment infrastructure and processes for multiple projects — Developing scripts for build, deployment, maintenance and related tasks using Jenkins, Docker, Maven, Golang and Bash — Experience in designing and deploying AWS Solutions using EC2, S3, and Elastic Load balancer (ELB), auto-scaling groups. — Used Micro services architecture with Spring Boot based service through REST —Creating Lambda function to automate snapshot back up on AWS and set up the scheduled backup. — Created and managed a Docker deployment pipeline for custom application images in the cloud using Jenkins.