Services
Video meeting . 15 mins
Priority DM . 2 days reply
Priority DM . 2 days reply
Video meeting . 30 mins
Video meeting . 30 mins
Video meeting . 30 mins
Priority DM . 2 days reply
Video meeting . 30 mins
Video meeting . 30 mins
Video meeting . 60 mins
About me
Data Engineering: Apache Spark, Databricks, AWS EMR, Glue, Azure Synapse, Data Warehousing, Data Modelling
Programming Languages: Python, SQL
Cloud Technologies: AWS, AZURE, Google Cloud
CI/CD: Jenkins, Azure Pipelines, Gitlab and AWS Code Deploy
Databases: MySQL, MongoDB, PL/SQL, Oracle 9i, 10g, SQL Server 2005, 2008, MS SQL, Postgres, Cassandra, Couchbase, Oracle NoSQL, DynamoDB.
IDE: VS Code, IntelliJ, Eclipse, Netbeans, SQL Developer, JDeveloper.
Designed and implemented self-healing ingestion pipelines to extract hourly data from Salesforce, Cisco, and SFTP sources, enhancing the Sales Customer Data Platform (CDP) reliability. Improved CDC framework in bash to ingest data from SAP views into HDFS using Cloudera HDP, catering to commodities clients.
Expertly curated data from the raw layer, applying Slowly Changing Dimension (SCD) Type 2 methods to update reporting tables with incremental data, ensuring accurate historical tracking.
Developed an intermediate data layer to streamline data processing, focusing on maintaining the latest records per unique ID at the ingestion step. This innovation reduced downstream report processing time by 40%.
Optimized PySpark code and ADF parallel copy activities, resulting in a 50% reduction in runtime for hourly data loads, significantly improving data pipeline efficiency.
Converted on-prem Teradata Data Warehouse reporting SQL scripts to Spark SQL, reducing reporting dataset creation by 60%. Migrated ingestion scripts to Databricks PySpark and leveraged Databricks features to cut down processing time.
Designed Spark-based ingestion pipelines for SCD Type 2 data from external insurance vendors using AWS EMR. Utilized CloudFormation to manage resources, create ephemeral EMR clusters, and define roles in IAM with minimal access following Data Governance standards. Built Step functions for converting Informatica flat files to parquet based on S3 event triggers and Lambda functions. Created Glue jobs for ad-hoc file processing within Step Functions, improving data availability.
Built Python scripts using Pandas and NumPy for validating dashboard data against data sources, ensuring data accuracy.