Services
Video meeting . 15 mins
Priority DM . 2 days reply
Video meeting . 30 mins
Video meeting . 30 mins
Video meeting . 30 mins
Priority DM . 2 days reply
Priority DM . 2 days reply
Video meeting . 30 mins
Video meeting . 30 mins
Video meeting . 60 mins
About me
• 8+ years of technical experience in Analysis, Design, Development with Big Data technologies like Spark, MapReduce, Hive, Kafka and HDFS including programming languages such as Python, Scala and Java in developing distributed data solutions, data analytical applications, and ETL pipelines by leveraging big data ecosystem components and AWS.
• Worked with Cloudera and Hortonworks distributions.
• Experience with Amazon EC2, Amazon S3, Amazon RDS, VPC, IAM, Amazon Elastic Load Balancing, Auto Scaling, Cloud Watch, SNS, SES, SQS, Lambda, EMR and other services of the AWS family.
• Worked on Amazon Web Services (AWS) Cloud Platform which includes services like EC2, S3, VPC, ELB, IAM, Dynamo DB, Cloud Front, Cloud Watch, Route 53, Elastic Beanstalk (EBS), Auto Scaling, Security Groups, EC2 Container Service (ECS), Code Commit, Code Pipeline, Code Build, Code Deploy, Dynamo DB, Auto Scaling, Security Groups, Red shift, Cloud Watch, Cloud Formation, Cloud Trail, Ops Works, Kinesis, IAM, SQS, SNS, SES.
• Excellent understanding of Hadoop architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming paradigm
• Experience with Apache Spark ecosystem Spark-Core, SQL, Data Frames, RDD's and knowledge on Spark MLLib.
• Extensive Knowledge on developing Spark Streaming jobs by developing RDD’s (Resilient Distributed Datasets) using Scala, PySpark and Spark-Shell.
• Extensive experience working on spark in performing ETL using Spark-SQL, Spark Core and Real-time data processing using Spark Streaming.
• Strong experience working with various file formats like Avro, Parquet, Orc, Json, Csv etc.
• Experience in developing customized UDF's in Python to extend Hive and Pig Latin functionality.
• Extensively worked with Teradata utilities Fast export, and Multi Load to export and load data to/from different source systems including flat files.
• Experienced in building Automation Regressing Scripts for validation of ETL process between multiple databases like Oracle, SQL Server, Hive, and Mongo DB using Python.
• Proficiency in SQL across several dialects (we commonly write MySQL, PostgreSQL, Redshift, SQL Server, and Oracle).