Services
Video meeting . 15 mins
Video meeting . 30 mins
Video meeting . 30 mins
Video meeting . 60 mins
Video meeting . 30 mins
Video meeting . 30 mins
Video meeting . 30 mins
About me
• 10 years of overall IT experience with 6 years in Datawarehouse & ETL and 3.5 years of experience in Bigdata software stacks such as HDFS, Sqoop, Hive, Sparksql, Spark, Kafka, Nifi, Oozie and Cloud platform of AWS (S3, EC2, Athena & EMR) and Azure(Azure Data Lake Storage,Azure Data Factory,Azure Databricks,Azure Synapse Analytics) .
• Experience in Banking and Healthcare domain.
TECHNICAL SKILLS
------------------------------
• Languages: Python
• ETL : IBM Datastage
• Bigdata ecosystem :
Distribution- Cloudera
Ingestion - Sqoop
SQL- Spark SQL, Hive
Processing – Spark
Scheduling- Oozie, Airflow
Web Interface- Hue
Streaming - Kafka
• Database : Oracle
• Reporting Tools – Tableau and Kibana
CERTIFICATIONS
------------------------------
• IBM certified solution developer Infosphere Datastage 8.0 and 8.5.
• Certified in Oracle SQL Expert
• Acquired ITIL foundation certificate in IT service management.
• Acquired domain certificate in BOA BFS -wealth management and Global consumer Banking
• Certified in Microsoft Azure data fundamentals
ROLES AND RESPONSIBITILY
----------------------------------
• Used Sqoop to efficiently transfer data between mysql database and HDFS with incremental load
• Worked on various file formats like Avro, Orc, Parquet
• Worked on file compression technique like Snappy
• Implemented partitioning and bucketing for best practices and
• improving performance of Hive queries.
• Involved in creating Hive-Hbase special table for faster access of customers lookup data
• Exposure to Kafka cluster, involved in connecting various Kafka Topics
• Involved in writing Scala code using Spark ,Spark SQL, Spark Streaming
• Involved in creating Jars and deploying code using spark-submit
• Exposure to Big data on cloud with AWS EMR,AWS Redshift, AWS Glue, AWS S3
PERFORMANCE OPTIMIZATION
------------------------------
• Used Broadcast Join, Partitioning & Bucketing on join columns
• Written spark Sql query in such a way to use Hash Aggregate over Sort Aggregate Used Repartitioning to increase parallelism
• Use Dataset Api to leverage Tungsten Optimization provided by Spark
• Use ReduceByKey over GroupByKey
• Used Salting on High volume data and keys with low cardinality in aggregation
• Used Cache or Persist after lot of transformations to save re calculation
• Use file formats like parquet with snappy compression for saving storage and faster processing
• Avoided or minimized the shuffling of data