Barcelona, Catalonia, Spain
As a seasoned Data Engineer with a career spanning over a decade, I thrive at the crossroads of data, technology, and strategic delivery. Across OLX, letgo, and most recently Adevinta, I have helped modernise data platforms—evolving legacy data warehouse and data management systems into robust, cloud-native and service-oriented architectures, with a strong focus on reliability, governance, operational ownership, and multi-user support. Areas of Expertise: - Data Lake & Lakehouse Solutions: Designing scalable foundations for analytics and decision-making, with clear raw/curated/serving patterns. - Data Ingestion & Streaming: Building robust batch and near-real-time pipelines with strong idempotency, data quality controls, and predictable operations. - Cloud Migration & Modernisation: Executing complex migrations (including vendor exits and global-to-domain de-globalisation), multi-year backfills, and production cutovers with minimal downtime. - DevOps, DataOps & MLOps Enablement: Applying CI/CD, infrastructure as code, and platform practices to run data systems in production; enabling ML use cases through reliable, well-governed data foundations. Throughout my career, I have blended the roles of Data Engineer and Cloud Data Architect, specialising in AWS-based solutions and modern data ecosystems (including Spark, Databricks, and Kafka). I bring hands-on experience building resilient data architectures, orchestrating large-scale data flows, and implementing pragmatic data governance—helping organisations move faster without compromising consistency, security, or trust in their data.
Biopharma R&D
Worked on Project Delorean, supporting the de-globalisation of legacy eBay analytics into Business Unit–owned data products, aligned with Adevinta’s decentralised data strategy. Primary focus on Subito (Automobile.it), with additional contributions to Benelux and Kleinanzeigen. - Led the redesign and migration of large-scale legacy data pipelines into independently operated, domain-owned data services, improving clarity of ownership, operational reliability, and long-term maintainability. - Contributed to the design and evolution of shared data platform services, including ingestion frameworks, schema validation mechanisms, data quality checks, and operational monitoring used by multiple teams. - Designed and operated lakehouse-based pipelines using Databricks, Spark (batch & streaming), Delta Lake and Parquet. - Built ingestion and transformation pipelines across AWS and GCP, including BigQuery and GA4 ingestion patterns into Databricks. - Implemented CDC and SCD (Type 1 & 2) merge strategies with strong schema validation and controlled schema evolution. - Orchestrated complex workflows using Apache Airflow, including incremental processing, backfills, and SLA-driven pipelines. - Implemented business logic using both Spark SQL and programmatic Spark APIs (Scala / PySpark), depending on complexity and performance needs. - Deployed and operated data workloads and microservices on Kubernetes using Helm and ArgoCD, and provisioned AWS infrastructure with Terraform / Terragrunt (S3, DynamoDB, SQS, cross-account setups). - Contributed to privacy-critical initiatives, building vendor-facing integrations for Data Deletion Requests (DDR) and applying GDPR-compliant data handling (anonymisation, access control, lineage).
Contributed to the Cloudera exit migration programme, replacing legacy ingestion and streaming workloads with a Confluent-based real-time platform while ensuring zero data loss and uninterrupted service. Acted as a senior technical reference, coordinating design decisions, supporting other engineers, and aligning platform changes across multiple teams and stakeholders. - Designed and operated real-time streaming pipelines using Apache Kafka, Kafka Streams, and Kafka Connect, including stateful enrichment with RocksDB. - Migrated legacy ingestion workloads (Cloudera / Flume) to Kafka-based architectures, ensuring zero data loss and uninterrupted service. - Worked on an operational data architecture, ingesting change events and writing enriched data into Oracle transactional databases, addressing concurrency and dual-datacenter consistency challenges. - Migrated and maintained streaming applications from Java 8 to Java 17 and Scala 2.13, implementing complex transformation logic using the Kafka Streams Scala API. - Supported CI/CD pipelines with GitHub Actions and adapted deployments to StatefulSets on OpenShift, collaborating closely with platform teams in a regulated banking environment. - Improved code quality and test coverage using scoverage and Sonar, and defined monitoring and alerting thresholds for Kafka lag and throughput.
Worked on the evolution of OLX’s analytics platform, focusing on performance, observability, and operational reliability of large-scale data warehouse workloads. - Applied DevOps and DataOps practices to data platforms, including CI/CD pipelines, infrastructure as code, containerised deployments, monitoring, and operational support for production data services. - Contributed to the design and rollout of a new global cloud-based data warehouse using AWS Redshift, supporting Autos and Global analytics use cases. - Built a comprehensive observability and alerting layer for ETL workloads, improving incident response times and operational visibility. - Analysed Redshift internal system tables to understand workload patterns, bottlenecks, and SLA breaches, and defined optimisation strategies. - Implemented monitoring, alerting, and on-call integrations using New Relic and PagerDuty. - Provided technical support and guidance to internal platform and support teams for configuring, operating, and troubleshooting data services. - Contributed Terraform modules to provision and manage Redshift infrastructure, monitoring components, and shared resources.
Designed and operated complex data migration and integration processes within shared data platforms, ensuring data consistency, service continuity, and controlled cutovers in production environments. - Built and operated monitoring and observability layers for data services, improving incident response, service reliability, and operational transparency for multiple stakeholder teams. - Designed and executed data migration pipelines using AWS DMS, Airflow, Redshift, and S3, ensuring data consistency and minimal downtime. - Rebuilt complex business logic using SQL-based transformations, validating volumes, duplicates, and integrity through automated checks. - Delivered curated datasets in Parquet on S3 for downstream consumption by the OLX Autos platform. - Provisioned and operated the full cloud infrastructure using Terraform, including DMS, Redshift, S3 buckets, and supporting resources. - Implemented data governance and compliance controls, retaining non-migrated datasets for legal and audit purposes.