Sandeep Mishra

Serving notice period | AWS Certified Data Platform Engineer @ Raymond James | S3 DataLake | Glue | Lambda | Spectrum | Dynamodb | Apache Spark Core Architecture | Yarn Architecture | Python | VPC| Redshift |AirFlow|

Bhubaneswar, Odisha, India

About

AWS Data Platform Engineer with 4.6 years of experience in building large-scale data pipelines, ETL/ELT processes, and data warehouse solutions. Utilized technologies like Python, SQL, Spark, Data lake, Databricks and Kafka to develop multi-terabyte scalable big data solutions for Fortune 500 Aviation and Retail companies. TECHNICAL SKILLS ----------------------- Programming Languages: Python, Scala, SQL Big Data Technologies: Spark, PySpark, Spark SQL, YARN, Hadoop, Hive, HBase Cloud Computing: AWS (Ec2,Emr,S3,Glue,Athena,Redshift,Lambda, Stepfunction), Azure (Databricks,ADF,ADLS) Data Engineering Tools: Data Modelling, ETL/ELT data Pipeline

Experience

  • AWS Data Engineer || ext @ Raymond James at Astrosoft Technologies
    Apr 2025 - Present · 1 yr 4 mos

    Oracle (On Prem) To Redshift (AWS Cloud) Migration Framework Kafka Streaming Medallion Framework (Migration + Transformation)

  • SplashBI ()
    • Associate Consultant
      Oct 2023 - Mar 2025 · 1 yr 6 mos

      - Migrated 40+ store databases from four zones to Amazon S3 using AWS DMS, ensuring scalable and centralized storage for downstream processing. - Utilized AWS Lambda to clean and process semi-structured JSON/CSV data, storing the refined data back into Amazon S3 for further transformation. - Crafted an ELT pipeline using AWS Glue (Crawlers, Jobs) and PySpark, transforming raw data into 10 Fact Tables and 40 Dimension Tables for efficient querying in Amazon Redshift. - Built optimized Data Marts in Redshift for various business needs, leveraging Amazon Athena, Redshift Spectrum, and AWS QuickSight, leading to a 15% increase in customer retail business profit through enhanced reporting and decision-making.

    • Associate Software Engineer
      Oct 2022 - Sep 2023 · 1 yr

      - Engineered and optimized data pipelines using dimension and fact tables for structured data storage and querying. Efficiently managed large datasets with daily incremental loads in Hive-style partitioned file formats for scalable and performant data retrieval. - Designed and implemented complex data transformations using AWS Glue Visual ELT, performing multi-stage joins, schema changes, filters to derive meaningful insights from flight, airport datasets and enhancing downstream analytics on Redshift tables. - Orchestrated end-to-end data workflows using AWS EventBridge and Step Functions to trigger Glue crawlers, execute ELT jobs, publish notifications for daily incremental data loads in S3 buckets and monitoring execution statuses using SNS and CloudWatch alerts. - Improved ELT performance by 15% reducing data processing latency with partition pruning and optimized Redshift COPY commands, and secure, cost-effective resource utilization.

    • Intern
      Apr 2022 - Sep 2022 · 6 mos

      - Learned the whole architecture of data cyclic movement in an end to end etl work flow processing also including data visualization tool.