Bhubaneswar, Odisha, India
AWS Data Platform Engineer with 4.6 years of experience in building large-scale data pipelines, ETL/ELT processes, and data warehouse solutions. Utilized technologies like Python, SQL, Spark, Data lake, Databricks and Kafka to develop multi-terabyte scalable big data solutions for Fortune 500 Aviation and Retail companies. TECHNICAL SKILLS ----------------------- Programming Languages: Python, Scala, SQL Big Data Technologies: Spark, PySpark, Spark SQL, YARN, Hadoop, Hive, HBase Cloud Computing: AWS (Ec2,Emr,S3,Glue,Athena,Redshift,Lambda, Stepfunction), Azure (Databricks,ADF,ADLS) Data Engineering Tools: Data Modelling, ETL/ELT data Pipeline
Oracle (On Prem) To Redshift (AWS Cloud) Migration Framework Kafka Streaming Medallion Framework (Migration + Transformation)
- Migrated 40+ store databases from four zones to Amazon S3 using AWS DMS, ensuring scalable and centralized storage for downstream processing. - Utilized AWS Lambda to clean and process semi-structured JSON/CSV data, storing the refined data back into Amazon S3 for further transformation. - Crafted an ELT pipeline using AWS Glue (Crawlers, Jobs) and PySpark, transforming raw data into 10 Fact Tables and 40 Dimension Tables for efficient querying in Amazon Redshift. - Built optimized Data Marts in Redshift for various business needs, leveraging Amazon Athena, Redshift Spectrum, and AWS QuickSight, leading to a 15% increase in customer retail business profit through enhanced reporting and decision-making.
- Engineered and optimized data pipelines using dimension and fact tables for structured data storage and querying. Efficiently managed large datasets with daily incremental loads in Hive-style partitioned file formats for scalable and performant data retrieval. - Designed and implemented complex data transformations using AWS Glue Visual ELT, performing multi-stage joins, schema changes, filters to derive meaningful insights from flight, airport datasets and enhancing downstream analytics on Redshift tables. - Orchestrated end-to-end data workflows using AWS EventBridge and Step Functions to trigger Glue crawlers, execute ELT jobs, publish notifications for daily incremental data loads in S3 buckets and monitoring execution statuses using SNS and CloudWatch alerts. - Improved ELT performance by 15% reducing data processing latency with partition pruning and optimized Redshift COPY commands, and secure, cost-effective resource utilization.
- Learned the whole architecture of data cyclic movement in an end to end etl work flow processing also including data visualization tool.