Arif Topkara

Cloud Data Engineer at ING Hubs Türkiye | 5x Google Cloud Certified

Istanbul, Türkiye

About

Cloud Data Engineer with 4+ years of experience designing and delivering scalable data solutions on Google Cloud. Currently working at ING Hubs Türkiye. I hold five Google Cloud Professional certifications — Data Engineer, ML Engineer, DevOps Engineer, Cloud Architect, and Database Engineer. My background spans both large-scale on-premise and cloud-native environments. At Migros, one of Turkey's largest e-commerce and retail companies, I worked across a hybrid architecture combining Hadoop (Spark, Hive, Impala, HDFS), on-premise Teradata, and Google Cloud services. I was responsible for developing and maintaining ETL pipelines, data modeling, Airflow DAG development, and contributing to the migration of Teradata and Hadoop workloads to Google Cloud — all within an agile, cross-functional team. Some of the key projects I have contributed to include: → Developing a BigQuery Agent using the Conversational Analytics API, deployed on Cloud Run and integrated into a client's internal application — with an automated metadata generation pipeline built on Dataplex, BigQuery data profiling and quality scans, and Terraform for infrastructure provisioning. → Building a configuration-driven ETL pipeline using Cloud Composer, Dataflow, Dataform, and BigQuery — enabling new data sources and tables to be onboarded through a single YAML metadata file, significantly reducing development overhead. → Architecting a Customer Data Platform by integrating data from Couchbase, SAP, GA4, and PostgreSQL into BigQuery, enabling customer segmentation for personalized experiences across SMS, email, campaigns, and mobile interfaces. → Developing real-time event processing pipelines using Pub/Sub, Dataflow, and Cloud Run to ingest and route web and mobile user event data into BigQuery. → Contributing to churn prediction and customer lifetime value projects using the RFM methodology. Core tech stack: Python, SQL, BigQuery, Cloud Composer (Airflow), Dataflow, Pub/Sub, Cloud Run, Dataform, Dataplex, Terraform, Apache Spark, Hive, Teradata, GCS. I graduated from Istanbul Technical University with a degree in Industrial Engineering, and during my studies served as Organization Coordinator at the ITU Data Science Society, where I helped run knowledge-sharing and workshop events. I enjoy tackling complex data problems and designing systems that are both technically sound and aligned with real business value.

Experience

  • Cloud Data Engineer at ING Hubs Türkiye
    May 2026 - Present · 3 mos

    I’m currently working on a project for ING Netherlands.

  • Cloud Data Engineer at NGC | Google Cloud Premier Partner
    Aug 2024 - Apr 2026 · 1 yr 9 mos

    • Designed, built, and optimized scalable data pipelines using Apache Beam (Dataflow) and Pub/Sub to ingest, transform, and load data into BigQuery • Developed a BigQuery Agent using Conversational Analytics API, deployed on Cloud Run and integrated into client’s internal application; built automated metadata generation pipeline with Dataplex, BigQuery profiling, and Terraform, reducing manual effort by 80% • Architected configuration-driven ETL pipeline using Cloud Composer, Dataflow, Dataform, and BigQuery; enabled new data sources via single YAML file, cutting development overhead by 70% • Architected Customer Data Platform integrating Couchbase, SAP, GA4, and PostgreSQL into BigQuery, enabling advanced customer segmentation and personalized campaigns across SMS, email, and mobile • Built real-time event processing pipelines with Pub/Sub, Dataflow, and Cloud Run, successfully handling 50,000+ events per second with sub-second latency • Implemented real-time monitoring and alerting using Pub/Sub, Dataflow, and Cloud Monitoring, reducing mean time to incident resolution by 60% • Collaborated with clients on optimal BigQuery data models (star schema, fact/dimension tables, partitioning & clustering), improving query performance by 35% • Identified and resolved performance bottlenecks in Dataflow pipelines and BigQuery SQL, delivering 20-25% cost savings on monthly GCP spend across client engagements • Led seamless on-premises to Google Cloud migrations using Dataflow and Data Fusion, migrating terabyte-scale data with zero downtime and achieving 30% reduction in operational costs • Provided consulting on data governance, security, and cost optimization best practices, resulting in average 20% reduction in cloud expenditures while maintaining compliance • Leveraged Data Fusion and Terraform to build reusable data integration patterns, accelerating project delivery by 50% and enhancing long-term maintainability

  • Migros Ticaret A.Ş. ()
    • Data Engineer
      Jan 2023 - Aug 2024 · 1 yr 8 mos

      • Engineered end-to-end data pipelines in a complex hybrid architecture (Teradata, Hadoop, GCP), ensuring 99.5% data availability for critical business units across Finance, E-commerce, and Marketing. • Led the migration of legacy Teradata and Hadoop workloads to Google Cloud, refactoring complex SQL queries and ETL processes that resulted in a 40% reduction in execution time for weekly CRM summaries. • Developed automated API-driven ingestion frameworks using Python and REST APIs (LiveNX, Microsoft Graph) to process store network and collaboration data, reducing manual data collection eforts by 60%. • Architected robust logical and physical data models using Star Schema methodologies; optimized fact/dimension structures in BigQuery that improved dashboard rendering speeds for cross-brand retail analytics by 30%. • Deployed a real-time security streaming pipeline using Pub/Sub and Dataflow to process RPA-generated data, enabling instant detection and alerting of information vulnerabilities across web platforms. • Optimized operational monitoring by establishing comprehensive alerting systems using Apache NiFi and Airflow, decreasing the mean time to resolution (MTTR) for pipeline failures by 35%. • Orchestrated complex workflows by developing and troubleshooting Airflow DAGs, ensuring seamless data movement across HDFS, Hive, and BigQuery for large-scale analytical workloads. • Collaborated in an Agile environment within cross-functional teams (Product Owners, Data Scientists, and Analysts) to translate complex business requirements into high-performance, scalable data solutions.

    • Data Warehouse and Business Intelligence Developer
      Nov 2021 - Dec 2022 · 1 yr 2 mos

      • Optimized complex SQL queries on Teradata to support high-priority business reporting, reducing data retrieval times for large-scale retail datasets and improving overall warehouse performance. • Designed and deployed executive-level interactive dashboards using MicroStrategy, providing Finance, Marketing, and Sales teams with real-time visibility into critical KPIs such as Like-for-Like (LFL) sales and gross margin. • Applied Data Warehousing principles (Kimball methodology) to design and maintain dimensional models (Star/Snowflake Schemas), ensuring high data integrity and consistent reporting across the enterprise. • Collaborated closely with cross-functional stakeholders in an Agile/Scrum environment to translate ambiguous business requirements into robust technical specifications and high-performance data views. • Conducted deep-dive data profiling and quality assessments to identify and resolve anomalies, ensuring a "single source of truth" and high user trust in self-service BI platforms. • Developed automated data validation scripts to monitor daily ETL loads, proactively identifying potential discrepancies before they impacted business-critical reports.

  • Organization and Sponsorship Coordinator at ITU Data Science Society
    Sep 2021 - Sep 2022 · 1 yr 1 mo

  • Data Analyst Intern at UPS
    Jun 2021 - Aug 2021 · 3 mos

    • Achieved good knowledge about SQL, Microsoft Access, Excel, Excel Solver and PowerBI • Gathered data from regional directorates for the Service Area Visiblity project and cleaned data using Excel