Greater Madrid Metropolitan Area
Staff Research Engineer at Rithum (Strategic Innovation Team). I build the ML systems that help one of the world's largest commerce networks understand its own data: product categorization and taxonomy unification, catalog-to-channel mapping, entity resolution across a brand/supplier network, GMV forecasting, return causality and attribution, and channel expansion recommendations. My work runs the full arc: research and model design, governed evaluation, production pipelines on lake-native infrastructure, and the design docs and strategy that align engineering and product around what the data actually says. I've set team standards for ML evaluation and iteration. I care as much about how we know a model works as about building it. I lead by making problems legible: quantifying gaps other teams can't see in their own pipelines, connecting evidence to a concrete ask, and convening the right people across engineering, data platform, and product to act on it. Within SIT I coordinate the entity resolution workstream, mentor engineers onboarding to our data platform, and help shape the team's technical direction. Background in computer vision and NLP research; left academia to solve high-impact problems at scale, and haven't looked back.
ML research and engineering for R/Intelligence, Rithum's commerce intelligence product suite, on the Strategic Innovation Team. ◦ Magic Mapper Categorizer: core ML behind Rithum's AI product categorization. Selected as the company-wide categorization engine after winning governed head-to-head evaluations; now unified across both major platform pipelines. Powered by RithumIQ ◦ R/Intelligence products: GMV forecasting, return causality & attribution, channel expansion recommendations, listing error detection, content optimization. ◦ Lead the entity resolution workstream (Splink) building a canonical brand/supplier network — coordinating engineers across the team and framing ER as a full ML product with tracked metrics and human feedback loops. ◦ Design and ship lake-native data pipelines (Databricks, Spark, dbt, Prefect); contribute to data catalog governance and table lifecycle standards. ◦ Set team standards for ML evaluation and iteration; run design reviews; mentor engineers onboarding to the data platform. Stack. Databricks, PySpark, Delta/Unity Catalog, dbt, Prefect, MLflow, Splink, CatBoost, VLLM, Qdrant, EconML.
ML engineering for AI-assisted product listing (Cadeera team). ◦ Built the canonical product categorizer and unified cross-channel taxonomy powering AI category suggestions. ◦ Categorization re-ranking models: trained, evaluated, and integrated into production APIs; scaled the system across channels. ◦ Designed Attribute2Field linking (AI suggestions between channel template fields and inventory attributes) through platform engineering design review. ◦ Established the team's ML strategy and iteration framework, and formal categorization benchmarks (accuracy, consistency, taxonomy adherence). Stack. Python, PyTorch, transformers, XGBoost, Qdrant, AWS ECS, Docker.
E-commerce: product aggregation. Won Most Innovative Tech by Hustle Awards. ◦ Image and textual-based product matching and variant detection. ◦ Product attribute extraction, colour detection, pose detection. ◦ Image processing pipelines. ◦ Entity parsing and named entity recognition. ◦ Product title construction. Stack. PyTorch, OpenCV, spaCy, xgboost, transformers, Faiss, Elasticsearch, AirFLow. GitHub workflows. Docker. MySQL. AWS: lambda, API Gateway, RDS, ECR.
Delivered real-time traffic analytics for one of Europe's largest highway operators. ◦ Vehicle detection, classification and tracking, speed estimation, lane detection. ◦ Deployment in Jetson Xavier on 3 different locations with 9 cameras. Stack. Python: MXNet, OpenCV, ONNX. GitLab CI/CD, Docker.
Also relocated to Madrid from Mexico. Of course, it coincided with the COVID pandemic.