Testimonials
Services
Video meeting . 60 mins
5
Optimizing spark jobs and understanding internals
Use spark resources to the fullest
Video meeting . 30 mins
5
Video meeting . 60 mins
Building Enterprise-Scale Data Products
Building Data Products the right way
Video meeting . 60 mins
4.5
Databricks Compute Cost Savings Review
Save cloud opex bill
Video meeting . 60 mins
5
Video meeting . 60 mins
Video meeting . 60 mins
4.7
Cloud Compute Cost Savings Review
Save cloud opex bill
Video meeting . 60 mins
5
Data engineering Architecture Review
Optimize the data pipelines
Video meeting . 60 mins
5
About me
Dynamic and results-driven versatile Senior Data Engineering Leader with over 15+ years of expertise in designing and implementing advanced data solutions. Proven track record in leading teams and managing large-scale data engineering projects. Expert in optimizing data pipelines, migrating legacy systems to modern technologies, and leveraging performance-enhancing techniques. Adept at driving strategic initiatives, ensuring data quality, and delivering actionable insights to support business objectives. Skilled in stakeholder management, cross-functional collaboration, and aligning data strategies with organizational goals. Committed to fostering innovation and excellence in data engineering practices.
AREA OF EXPERTISE
• Data Architecture & Strategy
• ML Engineering/ MlOps
• Team Leadership & Mentorship
• Agile & DevOps Methodologies
• Data Governance & Compliance
• Stakeholder Management & Communication
• Building Data Lakes, Data Fabric.
• Product Planning and OKR Definition.
• Leading MVP Definition and Development.
• Strategic Annual and Quarterly Planning.
• Project and Resource Planning.
• Optimization and simulation.
• Cross-functional team management.
• Data security, privacy, and Governance.
• Tactical Execution/ Project Oversight.
• Quantitative analysis
TECHNICAL SKILLS
• ML Engineering: MLflow, Kubeflow, Feature Store, Unity catalog, Model Registry.
• Hadoop & Spark Stack: Hadoop Ecosystem, Spark 2.x, Datafram API, Spark SQL, Databricks and Map Reduce, Alluxio.
• Streaming Stack: Spark Structure Streaming, Kafka, Kinesis, MSK, Flume.
• Database Technologies: Teradata, Oracle, Hive, Athena, Presto.
• Cloud: AWS, EKS, EC2, S3, Glue, EMR, Lambda, CloudWatch, SNS, SQS, EKS.
• No SQL: HBase, DynamoDB
• Languages: Scala, Python and Shell
• Monitoring and Reporting: Kubernetes, Airflow, Tableau, Looker and UC4
• Formats: Parquet, Delta, Iceberg.
• CI/CD: GIT, Jenkins, Nexus, Seldon Core.
• Security & Privacy: Privacera, Kerberos, Ranger, Audits, CPPA & GDPR data privacy.