Jaime Torres

Senior Data Engineer at AstraZeneca

Barcelona, Catalonia, Spain

About

As a seasoned Data Engineer with a career spanning over a decade, I thrive at the crossroads of data, technology, and strategic delivery. Across OLX, letgo, and most recently Adevinta, I have helped modernise data platforms—evolving legacy data warehouse and data management systems into robust, cloud-native and service-oriented architectures, with a strong focus on reliability, governance, operational ownership, and multi-user support. Areas of Expertise: - Data Lake & Lakehouse Solutions: Designing scalable foundations for analytics and decision-making, with clear raw/curated/serving patterns. - Data Ingestion & Streaming: Building robust batch and near-real-time pipelines with strong idempotency, data quality controls, and predictable operations. - Cloud Migration & Modernisation: Executing complex migrations (including vendor exits and global-to-domain de-globalisation), multi-year backfills, and production cutovers with minimal downtime. - DevOps, DataOps & MLOps Enablement: Applying CI/CD, infrastructure as code, and platform practices to run data systems in production; enabling ML use cases through reliable, well-governed data foundations. Throughout my career, I have blended the roles of Data Engineer and Cloud Data Architect, specialising in AWS-based solutions and modern data ecosystems (including Spark, Databricks, and Kafka). I bring hands-on experience building resilient data architectures, orchestrating large-scale data flows, and implementing pragmatic data governance—helping organisations move faster without compromising consistency, security, or trust in their data.

Experience

  • Senior Data Engineer at AstraZeneca
    May 2026 - Present · 3 mos

    Biopharma R&D

  • Data Platform Engineer at Adevinta Spain
    Apr 2025 - Apr 2026 · 1 yr 1 mo

    Worked on Project Delorean, supporting the de-globalisation of legacy eBay analytics into Business Unit–owned data products, aligned with Adevinta’s decentralised data strategy. Primary focus on Subito (Automobile.it), with additional contributions to Benelux and Kleinanzeigen. - Led the redesign and migration of large-scale legacy data pipelines into independently operated, domain-owned data services, improving clarity of ownership, operational reliability, and long-term maintainability. - Contributed to the design and evolution of shared data platform services, including ingestion frameworks, schema validation mechanisms, data quality checks, and operational monitoring used by multiple teams. - Designed and operated lakehouse-based pipelines using Databricks, Spark (batch & streaming), Delta Lake and Parquet. - Built ingestion and transformation pipelines across AWS and GCP, including BigQuery and GA4 ingestion patterns into Databricks. - Implemented CDC and SCD (Type 1 & 2) merge strategies with strong schema validation and controlled schema evolution. - Orchestrated complex workflows using Apache Airflow, including incremental processing, backfills, and SLA-driven pipelines. - Implemented business logic using both Spark SQL and programmatic Spark APIs (Scala / PySpark), depending on complexity and performance needs. - Deployed and operated data workloads and microservices on Kubernetes using Helm and ArgoCD, and provisioned AWS infrastructure with Terraform / Terragrunt (S3, DynamoDB, SQS, cross-account setups). - Contributed to privacy-critical initiatives, building vendor-facing integrations for Data Deletion Requests (DDR) and applying GDPR-compliant data handling (anonymisation, access control, lineage).

  • Senior Data Engineer at Santander UK
    Jan 2024 - Mar 2025 · 1 yr 3 mos

    Contributed to the Cloudera exit migration programme, replacing legacy ingestion and streaming workloads with a Confluent-based real-time platform while ensuring zero data loss and uninterrupted service. Acted as a senior technical reference, coordinating design decisions, supporting other engineers, and aligning platform changes across multiple teams and stakeholders. - Designed and operated real-time streaming pipelines using Apache Kafka, Kafka Streams, and Kafka Connect, including stateful enrichment with RocksDB. - Migrated legacy ingestion workloads (Cloudera / Flume) to Kafka-based architectures, ensuring zero data loss and uninterrupted service. - Worked on an operational data architecture, ingesting change events and writing enriched data into Oracle transactional databases, addressing concurrency and dual-datacenter consistency challenges. - Migrated and maintained streaming applications from Java 8 to Java 17 and Scala 2.13, implementing complex transformation logic using the Kafka Streams Scala API. - Supported CI/CD pipelines with GitHub Actions and adapted deployments to StatefulSets on OpenShift, collaborating closely with platform teams in a regulated banking environment. - Improved code quality and test coverage using scoverage and Sonar, and defined monitoring and alerting thresholds for Kafka lag and throughput.

  • Data Platform Engineer at OLX
    Nov 2022 - Dec 2023 · 1 yr 2 mos

    Worked on the evolution of OLX’s analytics platform, focusing on performance, observability, and operational reliability of large-scale data warehouse workloads. - Applied DevOps and DataOps practices to data platforms, including CI/CD pipelines, infrastructure as code, containerised deployments, monitoring, and operational support for production data services. - Contributed to the design and rollout of a new global cloud-based data warehouse using AWS Redshift, supporting Autos and Global analytics use cases. - Built a comprehensive observability and alerting layer for ETL workloads, improving incident response times and operational visibility. - Analysed Redshift internal system tables to understand workload patterns, bottlenecks, and SLA breaches, and defined optimisation strategies. - Implemented monitoring, alerting, and on-call integrations using New Relic and PagerDuty. - Provided technical support and guidance to internal platform and support teams for configuring, operating, and troubleshooting data services. - Contributed Terraform modules to provision and manage Redshift infrastructure, monitoring components, and shared resources.

  • Data Platform Engineer at OLX Autos
    Jan 2022 - Nov 2022 · 11 mos

    Designed and operated complex data migration and integration processes within shared data platforms, ensuring data consistency, service continuity, and controlled cutovers in production environments. - Built and operated monitoring and observability layers for data services, improving incident response, service reliability, and operational transparency for multiple stakeholder teams. - Designed and executed data migration pipelines using AWS DMS, Airflow, Redshift, and S3, ensuring data consistency and minimal downtime. - Rebuilt complex business logic using SQL-based transformations, validating volumes, duplicates, and integrity through automated checks. - Delivered curated datasets in Parquet on S3 for downstream consumption by the OLX Autos platform. - Provisioned and operated the full cloud infrastructure using Terraform, including DMS, Redshift, S3 buckets, and supporting resources. - Implemented data governance and compliance controls, retaining non-migrated datasets for legal and audit purposes.