Services

Video meeting . 15 mins
FREE
Video meeting . 30 mins
800
Video meeting . 30 mins
800
Video meeting . 60 mins
1,300
Video meeting . 30 mins
800
Video meeting . 30 mins
800
Video meeting . 30 mins
800
Popular

About me

• 10 years of overall IT experience with 6 years in Datawarehouse & ETL and 3.5 years of experience in Bigdata software stacks such as HDFS, Sqoop, Hive, Sparksql, Spark, Kafka, Nifi, Oozie and Cloud platform of AWS (S3, EC2, Athena & EMR) and Azure(Azure Data Lake Storage,Azure Data Factory,Azure Databricks,Azure Synapse Analytics) . • Experience in Banking and Healthcare domain. TECHNICAL SKILLS ------------------------------ • Languages: Python • ETL : IBM Datastage • Bigdata ecosystem : Distribution- Cloudera Ingestion - Sqoop SQL- Spark SQL, Hive Processing – Spark Scheduling- Oozie, Airflow Web Interface- Hue Streaming - Kafka • Database : Oracle • Reporting Tools – Tableau and Kibana CERTIFICATIONS ------------------------------ • IBM certified solution developer Infosphere Datastage 8.0 and 8.5. • Certified in Oracle SQL Expert • Acquired ITIL foundation certificate in IT service management. • Acquired domain certificate in BOA BFS -wealth management and Global consumer Banking • Certified in Microsoft Azure data fundamentals ROLES AND RESPONSIBITILY ---------------------------------- • Used Sqoop to efficiently transfer data between mysql database and HDFS with incremental load • Worked on various file formats like Avro, Orc, Parquet • Worked on file compression technique like Snappy • Implemented partitioning and bucketing for best practices and • improving performance of Hive queries. • Involved in creating Hive-Hbase special table for faster access of customers lookup data • Exposure to Kafka cluster, involved in connecting various Kafka Topics • Involved in writing Scala code using Spark ,Spark SQL, Spark Streaming • Involved in creating Jars and deploying code using spark-submit • Exposure to Big data on cloud with AWS EMR,AWS Redshift, AWS Glue, AWS S3 PERFORMANCE OPTIMIZATION ------------------------------ • Used Broadcast Join, Partitioning & Bucketing on join columns • Written spark Sql query in such a way to use Hash Aggregate over Sort Aggregate Used Repartitioning to increase parallelism • Use Dataset Api to leverage Tungsten Optimization provided by Spark • Use ReduceByKey over GroupByKey • Used Salting on High volume data and keys with low cardinality in aggregation • Used Cache or Persist after lot of transformations to save re calculation • Use file formats like parquet with snappy compression for saving storage and faster processing • Avoided or minimized the shuffling of data