Services
Video meeting . 30 mins
Video meeting . 30 mins
Video meeting . 30 mins
Quick chat( No Technical Stuffs)
Quick chat( No Technical Stuffs)
Priority DM . 2 days reply
Video meeting . 30 mins
Video meeting . 30 mins
Priority DM . 2 days reply
Priority DM . 2 days reply
Priority DM for Short Queries
Priority DM for Short Queries
Video meeting . 15 mins
Video meeting . 45 mins
About me
Data Engineer / Big Data Developer with 5+ years of experience in designing and implementing efficient ETL processes, resulting in significant project cost savings.
Specialized in Apache Spark, Pyspark, Python, SQL, Azure Databricks, Azure Datafactory, and Azure Data Lake.Proficient in Data Engineering, Data Pipelines, Data Modelling, Data warehousing, and performance optimization Worked on Batch pipeline and contains knowledge of Streaming pipelines also.
Professional Journey :
----------------------
1)Wipro Technologies:
I began my career at Wipro Technologies, developing Big Data solutions for the Retail sector, and managing the ingestion and processing of daily sales data using PySpark and Hadoop. Developed an end-to-end big data pipeline for a global retail store, achieving a 30% time saving through Spark SQL optimization.
2)Capgemini India:
Worked as a Data Engineer, leading the development of an entity management tool for deduplication of customer and policy records. Optimized data merging and enrichment using join optimization, improving performance by 10%. Earned praise from ADNIC CIO and stakeholders.
3)Michelin India:
Currently working as Deputy Manager - Data Engineer, developing Spark-based applications using Databricks for data extraction, transformation, and aggregation. Collaborated with data scientists to integrate ML models with ADF batch pipelines, reducing processing time by 20%. Led performance tuning on SQL queries, optimizing data retrieval by 20%, and migrated projects from ADLS Gen1 to Gen2, saving 30% on storage costs.
Technical Skills :
-----------------
• Big Data Tools: Spark, Pyspark, Performance Tuning, Spark SQL, Hadoop, Hive, Spark Streaming
• Languages: SQL, Python
• Azure Services: Azure Data Lake Gen2, Azure Blob,(ADF),Azure Databricks(ADB)
• Databricks Specific: Delta Lake , Delta Tables, Unity Catalog, Autoloader
• AWS Services: Redshift, S3, Athena, EMR, Glue
• Databases: SQL, Advance SQL, MySql, MS SQL Server SSIS, Azure SQL
• Datawarehouse: Snowflake Cloud Data Warehouse
• Version Control: Git, Gitlab, Github, Azure Devops
• Methodologies & Tools: Jira, Jenkins, Pycharm, Vscode, pytest, Agile, Scrum, CI/CD
• Operating System: Windows, Linux
•Soft Skills: Communication, Teamwork, Leadership, Analytical Skills.
Feel free to connect with me via email at siddharthnahata2396@gmail.com
Keep Learning, Keep Growing.