Services
Priority DM . 2 days reply
Priority DM . 2 days reply
Video meeting . 30 mins
Video meeting . 30 mins
Video meeting . 30 mins
Priority DM . 2 days reply
Video meeting . 30 mins
Video meeting . 30 mins
Discovery Call
Have you a use case/opportunity for discovery? Let's connect
Video meeting . 60 mins
About me
A seasoned Lead Solutions Architect specializing in Apache Spark, Databricks, Snowflakes, Click House, Redshift, Synapse Analytics, Teradata & Netezza with experience in designing and implementing data solutions for big data warehouses and Lakehouse. Expertise in Spark’s runtime internals, query optimization, performance tuning and troubleshooting. Successfully migrated large-scale data warehouses to the Spark ecosystem and Delta Lake, to major data warehouse ecosystem databases, enhancing data architectures through Spark SQL optimization for performance and scalability. Proven ability to lead cross-functional teams and deliver innovative data platforms that meet business goals.
• Technical Expertise:
o Extensive experience with massive parallel processing database systems such as Databricks, Synapse Analytics, Redshift, Snowflake, Netezza, Teradata, and Click house.
o Proficient with transactional database systems including Postgres, MySQL, SQL Server, and MongoDB.
o Hands-on experience with data visualization tools like Mongo Charts and Power BI.
o Experienced in Delta Lake tables, Unity Catalog, spark streaming and ML capabilities
o Hands-on working experience on Apache Iceberg format, Apache Druid
o Normalization and De-normalization approaches
o Determining functions for Warehouse and building snowflake and star schema architectures for Slowly changing dimensions (1/2/3)
o Building data marts for subject oriented analytics and insight
o Data Governance prospects implementing through centralized catalog controls.
o Implementing decision support systems, Dimensional structures depicting roadmap for large DW/BI implementation.
o Expertise in building a unified Data Lakehouse platform using Databricks, leveraging its product catalog, data lineage, data governance, and pipeline orchestration services.
o Using GenAI for reference queries and understanding standard practices for DW/BI perspectives
o Experienced in designing and implementing real-time data pipelines using Apache Kafka and open-source solutions like Apache Druid, Click House, Apache Flink, and DBT.
o Building a Unified Data Governance solution using Unity Catalog to manage access privileges across clusters, users, and apps.
o Working on LLM, creating SLM, implementing RAG using vector Database, document cleaning, chunking and chunk embeddings
o Fine tuning open source LLM, few shot learning, Reinforcement learning
o Making sure the security and ethical aspects of Prompt injection, jail breaking, leaking and hallucinations are taken care