Services

Priority DM . 2 days reply
FREE
Priority DM . 2 days reply
FREE
Video meeting . 30 mins
$1
Popular
Priority DM . 2 days reply
FREE
Video meeting . 30 mins
$1
Video meeting . 30 mins
5
$1

Ratings and feedback

5/5
1 ratings

About me

A highly motivated and technically driven analytical professional with nearly 10 years of experience in Big Data and Machine Learning with expertise in designing data-intensive applications using the Hadoop Ecosystem, Big Data Analytics, Cloud Data Engineering, Data Structure, ML Algorithms, Data Warehouse, Data Visualization, and Data Quality. Imported the data from various formats like Mainframe, JSON, XML, Text, and CSV to HDFS cluster with compressed for optimization. Extensive Knowledge of Bigdata Enterprise architecture (Cloudera preferred). Designed and developed services to persist and read data from Hadoop, HDFS, and Hive, and wrote Java-based MapReduce batch jobs using Cloudera Hadoop Data Platform. Created Hive External tables and loaded the data into tables and query data using HQL. Using Hive join queries to join multiple tables of a source system and load tables. Implemented Python using PySpark SQL for faster testing and processing of data. Experience in tuning SQL queries to maximize performance. Migrated the computational code in HQL to PySpark. Experienced in scripting(Unix/Linux) and scheduling. Completed data extraction, aggregation, and analysis in HDFS by using PySpark. Developed Pre-processing job using Spark Data frames to flatten JSON documents to flat files. Strong understanding of distributed computing principles and experience with large-scale data processing frameworks. Creating mount points for cloud storage in DBFS to implement RBAC for end users. Experience in building frameworks for data ingestion, processing, and consumption using GCP Data Flow, GCP Data Composer, and BigQuery. Building frameworks for data ingestion, processing, and consumption using GCP Data Flow, GCP Data Composer, and BigQuery. Experience in developing and deploying data pipelines in GCP. Implement best practices for data management, security, and governance within the Databricks environment. Experience in Snowflake, BigQuery and/or Databricks experience. Experience of Databricks Unity Catalog and GitHub and CI/CD Pipelines. Experience in Airflow for Job Orchestration, dependency Setup and Job Scheduling. In-depth understanding of data modelling, data warehousing, and data integration concepts and best practices. Design, develop, and deploy Databricks jobs to process and analyze large volumes of data. Collaborate with data engineers and data scientists to understand data requirements and implement appropriate data processing pipelines. Optimize Databricks jobs for performance and scalability to handle big data workloads.