Reynolds Pravindev

Solutions Architect @ Databricks

Vancouver, British Columbia, Canada

About

Experience

  • Solutions Architect at Databricks
    Mar 2025 - Present · 1 yr 5 mos

  • Lead Data Engineer at Cognizant
    Mar 2022 - Mar 2025 · 3 yrs 1 mo

    Client - Albertsons Group Responsibilities: As a lead data engineer design, deliver and sustain next generation eCommerce workload planning automation and forecasting solution which helped achieve targeted labor costs and propel the sales by ~15%. • Built the solution architecture and development of an object-oriented, automated Workload Labor Automation and Forecasting platform with pyspark on Azure Databricks + Azure Data Factory which had integrations to Apache Kafka and Delta lake. The data was consumed by backend APIs built on Java 17 with feedback to database from UI. • Implemented a lakehouse architecture in conjunction with Azure Data Lake as a delta lake filesystem on Databricks. • Developed a pyspark-based streaming solution for reading and writing streaming feedback data to Apache Kafka. • Worked on a generative AI solution for auto search term classification using Azure Open AI models (Python SDK) using a combination of prompt engineering and embeddings search by using an interim vector DB for subsequent searches. • Revamped the logging and telemetry by leveraging Azure Log Analytics and Kusto(ADX) integration for Spark logs. • Worked on a POC for container-based model using Docker and Azure Container services (ACR). Rearchitected the legacy file upload process from FTP server to Azure Cosmos DB via pyspark. • Developed a metadata-based, reusable, pyspark library for common data load from various database sources like Snowflake, Azure SQL DB, Azure Data Lake and Apache Kafka. • Developed a solution for reading Delta tables offline without using Databricks clusters as a cost-saving initiative. • Worked on Spark performance optimization, data governance and protection i.e., GDPR and SCHREMS. Designed a disaster recovery model for the entire framework. Awards: Awarded for bringing back the customer trust by delivering robust data solutions with low defect ratio as well as being time efficient.

  • Lead Data Engineer at Electronic Arts (EA)
    Jan 2022 - Mar 2022 · 3 mos

    Responsibilities: As a lead data engineer, build a data mart of game user metrics data for "FIFA Mobile" which feeds data to reporting solution. The reports are then used by game analysts to model, predict and analyze user experience. • Worked on designing and developing ETL pipelines using Apache Airflow which includes Python programming. • Modelled and developed DB solutions on Snowflake using SnowSQL, Amazon Redshift • Developed Spark (pyspark) workloads on Databricks for game analytics. • Implemented data storage using AWS S3 and Azure Data Lake and Identity management using Azure AD. Cl/CD implementation via docker for container-based deployment, GitLab and Git. • Performance optimization on Redshift data warehouse queries

  • Infosys (6 yrs 8 mos)
    • Technology Lead
      Nov 2018 - Dec 2021 · 3 yrs 2 mos

      Client - Microsoft Corporation Responsibilities: As a lead data engineer, lead a team of 10 data engineers and drive various engineering projects. • Designed and built a real-time big data framework processing -6 TB data on daily basis using Azure Databricks with Delta Lake. Developed PowerShell wrapper with REST API to perform SPARK on Synapse operations using REST API and programmatically submitting Spark Session queries and Spark Batch jobs. • Worked on a generic python-based data validation framework which was completely metadata-driven to validate data between ADB and Synapse based workload executions. • Changes to the ASP .NET application to accommodate SPARK features on Synapse Analytics and also ensure feature parity with Azure Databricks. • Setup an external Hive metastore to share spark tables metadata between different compute platforms. • Performance studies comparing ADB and Synapse and outlining where the bottleneck was and addressing them as needed. • Designed and developed a completely metadata driven framework for ETL and data processing using Azure Synapse Spark, Python, pyspark and Azure Data Factory and built custom python wheel files for modularized code. • Designed and developed Cl/CD framework for automated infrastructure and code deployment using Azure DevOps, PowerShell, Azure PowerShell (ARM and Az), Azure CLI, Unix Shell scripting and Git shell. • Developed an alerting and monitoring frameworks using Azure monitor, Log Analytics with Azure Diagnostics logs and Azure Logic Apps for immediate response on failures, service outages and critical issues. • Build data privacy solutions for existing data using features like Dynamic Data Masking, Transparent Data Encryption, Code Analyzers, PII Scrubbing using python libraries like "scrubadub". Awards: Multiple awards for innovation and deliver excellence in the year 2020 and 2021

    • Technology Analyst
      May 2015 - Nov 2018 · 3 yrs 7 mos

      Client - Microsoft Corporation

  • Senior Software Engineer at Capgemini
    Feb 2012 - Apr 2015 · 3 yrs 3 mos

    Formerly IGATE. Client - Royal Bank of Canada Responsibilities: As a data engineer, develop and sustain an existing Live and Health Insurance policy administration system. • Involved in creating simple and complex SSIS packages to extract, transform and merge data from different database sources and use the resultant data in business logic implementation. • Involved in development of T-SQL stored procedures to address issues faced by business user in the application. Involved in Performance Tuning of T-SQL queries to improve database performance and availability. • Hands in experience in creating packages with complex transforms to extract data from external sources. Involved in preparing Technical Design Documents and Software Requirement Specifications document.