Services

Video meeting . 30 mins
750
Video meeting . 30 mins
1,500
Video meeting . 30 mins

Quick chat( No Technical Stuffs)

Quick chat( No Technical Stuffs)
500
Priority DM . 2 days reply
250
Video meeting . 30 mins
1,300
Popular
Video meeting . 30 mins
1,000
Priority DM . 2 days reply
250
Popular
Priority DM . 2 days reply

Priority DM for Short Queries

Priority DM for Short Queries
250
Video meeting . 15 mins
250
Video meeting . 45 mins
2,000

About me

Data Engineer / Big Data Developer with 5+ years of experience in designing and implementing efficient ETL processes, resulting in significant project cost savings. Specialized in Apache Spark, Pyspark, Python, SQL, Azure Databricks, Azure Datafactory, and Azure Data Lake.Proficient in Data Engineering, Data Pipelines, Data Modelling, Data warehousing, and performance optimization Worked on Batch pipeline and contains knowledge of Streaming pipelines also. Professional Journey : ---------------------- 1)Wipro Technologies: I began my career at Wipro Technologies, developing Big Data solutions for the Retail sector, and managing the ingestion and processing of daily sales data using PySpark and Hadoop. Developed an end-to-end big data pipeline for a global retail store, achieving a 30% time saving through Spark SQL optimization. 2)Capgemini India: Worked as a Data Engineer, leading the development of an entity management tool for deduplication of customer and policy records. Optimized data merging and enrichment using join optimization, improving performance by 10%. Earned praise from ADNIC CIO and stakeholders. 3)Michelin India: Currently working as Deputy Manager - Data Engineer, developing Spark-based applications using Databricks for data extraction, transformation, and aggregation. Collaborated with data scientists to integrate ML models with ADF batch pipelines, reducing processing time by 20%. Led performance tuning on SQL queries, optimizing data retrieval by 20%, and migrated projects from ADLS Gen1 to Gen2, saving 30% on storage costs. Technical Skills : ----------------- • Big Data Tools: Spark, Pyspark, Performance Tuning, Spark SQL, Hadoop, Hive, Spark Streaming • Languages: SQL, Python • Azure Services: Azure Data Lake Gen2, Azure Blob,(ADF),Azure Databricks(ADB) • Databricks Specific: Delta Lake , Delta Tables, Unity Catalog, Autoloader • AWS Services: Redshift, S3, Athena, EMR, Glue • Databases: SQL, Advance SQL, MySql, MS SQL Server SSIS, Azure SQL • Datawarehouse: Snowflake Cloud Data Warehouse • Version Control: Git, Gitlab, Github, Azure Devops • Methodologies & Tools: Jira, Jenkins, Pycharm, Vscode, pytest, Agile, Scrum, CI/CD • Operating System: Windows, Linux •Soft Skills: Communication, Teamwork, Leadership, Analytical Skills. Feel free to connect with me via email at siddharthnahata2396@gmail.com Keep Learning, Keep Growing.