Barcelona, Catalonia, Spain
I am a data enthusiast with a background in Bioinformatics currently working at the intersection between Data Science and Engineering. In my current role, I help deliver analytics-ready data to accelerate Data Science workflows in AstraZeneca R&D. My day to day is varied, performing tasks like data transfers, anonymisation of datasets, enforcing data quality or developing internal libraries for data/ML tasks in R and Python.
As a member of the ADAPT (Analytics DAta PreparaTion) team, I aid in the acceleration of Data Science analysis within the R&D Data Office. We help deliver analytics-ready data to reduce the burden of data wrangling in Data Science workflows, either by bespoke data preparation tasks, or by developing re-usable tools. Some of the projects I've contributed to while in this role include: - Extending an internal library for synthetic data generation using GANs, as well as researching and implementing evaluation metrics - Improving the performance of a synthetic data generation workflow by orchestrating parallel execution in an HPC cluster, as well as developing a bash CLI for the tool - Building a reproducible R pipeline to perform extensive Quality Control of a large RNAseq dataset and summarise the findings in an RMarkdown report - Ingesting a large public MRI database (ADNI) into an HPC environment and building a scalable data pre-processing pipeline using Singularity containers - Harmonising, curating and cleaning heterogenous data sources to ingest them into a Knowledge Graph using R and Python - Creating a customer-facing Streamlit dashboard as a GUI for requesting synthetic data - Developing an internal Python package for the anonimisation of clinical trial data - Productionising a Machine Learning methodology into a package while enforcing Software Development good practices
Project title: Transforming interpretation of metabolomics data by improving inegrative pathway analysis tools. Metabolomics is the field that studies the abundance of small molecules in biological samples and its association to health and disease. In recent years, it has benefited greatly from a family of methods termed Pathway Analysis: methods that rely on prior biological knowledge to abstract the measurement of many variables (metabolites) into one (pathways). During this 6-month project, I assessed the state of the art in metabolomics pathway analysis, and compared it to transcriptomics in terms of maturity and field-specific limitations. I then critically evaluated two popular Pathway Analysis methods: GSEA and Globaltest, by analysing how varying their parameters affected the downstream analysis. Finally, leveraging semi-synthetic data simulations, I explored the limitations of both methods under challenges like metabolite mis-identification, and concluded that Globaltest suffered from a high tendency to produce false positives; whereas GSEA provided more FP control.
Deep Learning approaches to enhance mass spectrometry hyperspectral images: towards spatial and spectral super resolution As a Postgraduate Research Associate, I carried out a project investigating the potential of using Deep Learning methods to produce super-resolved Mass Spectrometry Images (MSI). The project involved a phase of literature research, where I gained familiarity with the state of the art in image super-resolution. Later on I implemented and adapted two of the best-performing Deep Learning models and set up experiments to evaluate critically their performance. Finally, I analysed the data from said experiments looking at a variety of image quality metrics (from classical approaches like the Signal-to-Noise ratio, to more sophisticated ones based on DL), and concluded proposing strategies to improve the development of DL algorithms for super-resolution of Mass Spectrometry Images.
Tutoring second year students of BSc Biotechnology during a project analyzing RNA microarray expression data in R. The students learned hands-on how to manipulate and analyse large biological datasets in a reproducible manner. They learned and applied data-related and statistical analysis skills like hypothesis testing, dimensionality reduction, data visualisation, etc. I organised weekly meetings with the students where we would assess their progress, troubleshoot any technical issues and discuss questions related to the analysis plan.
I analyzed RNA-Seq data to inspect the effect of 3' Untranslated Regions on gene expression in Saccharomyces cerevisiae.
During this internship, I carried out two main tasks: I performed a biochemical assay to test DNA probes as diagnostic tools for bacterial infections. The experiment mainly involved measuring fluorescence of dozens of candidate probes in the presence of different pathogens. Additionally, and out of own intiative, I automated the data analysis pipeline (wrangling, visualisation and statistical analysis) using R, saving hours of manual and error-prone effort.