Gurugram, Haryana, India
Data Engineer with 2.3 years of professional Data Engineering experience, specializing in designing and building scalable ETL/ELT pipelines, CDC architectures, and AI-ready data platforms in the Telecom domain. Proven track record of processing 150M+ records across Oracle, MongoDB, MySQL, and AWS-based data platforms, delivering trusted datasets for analytics, business intelligence, and downstream AI/ML workloads. Currently working at Tecnotree, delivering enterprise-scale telecom data engineering, migration, and transformation solutions for MTN Cameroon and STC Kuwait. π― Core Impact β’ Processed 150M+ records across enterprise telecom data engineering projects for MTN Cameroon and STC Kuwait. β’ Built scalable PySpark pipelines on Databricks and Delta Lake using Medallion Architecture (Bronze β Silver β Gold) and watermark-based CDC. β’ Reduced distributed PySpark job execution time by ~40% using broadcast joins, partition pruning, and skew handling. β’ Improved analytical query performance by 2Γ through Star Schema design and SQL optimization. β’ Developed metadata-driven validation and reconciliation frameworks, reducing manual effort by ~60% while ensuring trusted production datasets. β’ Delivered enterprise-scale telecom data platforms supporting successful production go-lives in Cameroon and Kuwait. π Core Technologies Big Data & Cloud: PySpark, Spark SQL, Databricks, Delta Lake, Medallion Architecture, AWS (S3, RDS, Redshift, MWAA) Databases: Oracle, PostgreSQL, MySQL, MongoDB, Elasticsearch Data Engineering: ETL/ELT Pipelines, Watermark-based CDC, Apache Airflow (MWAA), Data Validation, Reconciliation, Docker, GitHub Actions Additional Experience: Azure Databricks, Azure Data Factory (ADF), ADLS Gen2, Azure Synapse (Personal Project) Passionate about building scalable, reliable, and high-performance data platforms that accelerate analytics, business intelligence, and data-driven decision making. π Immediate Joiner | Open to Data Engineer opportunities
Scale & Ingestion:- Engineered high-integrity ingestion workflows for 90M+ records from client-provided flat files stored in Amazon S3 into MongoDB using full-load and watermark-based CDC, supporting 2M+ active telecom subscribers with comprehensive validation and zero unresolved discrepancies. Distributed Processing:- Built scalable PySpark transformation pipelines on Databricks and Delta Lake using Medallion Architecture (Bronze β Silver β Gold), applying partition pruning, broadcast joins, and skew handling to reduce distributed processing time by ~40%. Data Modeling & Analytics:- Implemented Star Schema models (Fact: Subscriptions, Recharge; Dimensions: Customer, Plan, Date) in collaboration with data analysts and optimized analytical SQL queries, improving reporting performance by 2Γ and enabling enterprise reporting and predictive analytics. Orchestration & DevOps:- Developed and orchestrated fault-tolerant ETL workflows using Apache Airflow (MWAA), authored Python automation scripts for data processing and validation, containerized components with Docker, and integrated CI/CD using GitHub Actions. Data Reliability & Quality: Built metadata-driven validation and PostgreSQL-based reconciliation frameworks with Grafana dashboards, reducing manual reconciliation effort by ~60% and delivering trusted, analytics-ready datasets for downstream AI/ML workloads.
Worked on βData Routing Techniques for IoT Enabled Wireless Sensor Networks (WSNs)β at IIT (BHU), focusing on energy-efficient routing and clustering techniques to optimize sensor network performance. Implemented and analyzed multi-hop intra-clustering approaches using Python to reduce energy consumption and improve network lifetime in WSN environments. Studied routing optimization, cluster head selection, and hotspot reduction techniques for IoT-enabled sensor networks. Performed data analysis, simulation, and performance evaluation of wireless communication models, gaining hands-on exposure to Python programming, network simulation, research methodologies, and performance analysis.