Services

Video meeting . 45 mins
499999
Popular

About me

Experienced Data Engineer with over 10 years in software development, including more than 7 years focused on data platforms. Experience Highlights: • Built scalable data ingestion and processing frameworks using Python, Pandas, PySpark and Azure Databricks, incorporating data quality checks and efficient transfer of data from files to ADLS based on dynamic configurations. • Strong experience scheduling and orchestrating jobs with Azure Data Factory and Databricks Workflows. • Built end-to-end data pipelines in Microsoft Fabric – Synapse Data Engineering using Fabric Pipelines for orchestration and Fabric Notebooks for transformation, leveraging Lakehouse architecture and OneLake for unified storage. • Experienced in batch and real-time streaming data processing using Azure Event Hubs and Spark Structured Streaming for near real-time ingestion and analytics. • Hands-on experience in Spark performance optimization, analysing scan, shuffle, spill, skew, and data serialization metrics to improve job execution. • Skilled in estimating and configuring Databricks clusters to meet workload demands and SLA requirements. • Proficient in working with multiple file formats, including JSON, CSV, and columnar formats like Parquet and Delta. • Strong experience in working with MS SQL Server and Oracle databases - writing complex SQL queries, stored procedures, triggers, views, and cursors.