Testimonials

Services

Digital Product
5

Databricks - Big Book of Data Engineering

databricks book with 5+ case studies
₹0₹200
Popular
Priority DM . 2 days reply
5
₹50₹249
Popular
Digital Product
4.8

Master Python with this notes

One book you need to master Python
FREE
Priority DM . 3 days reply

DM for referral

Share your updated Resume with me
₹50
Video meeting . 40 mins

Resume Review

Let's get it sorted
₹599₹1,000
Priority DM . 3 days reply
₹50
Video meeting . 35 mins
5

1:1 Mentorship (personalized)

20+ LPA is possible, you need some guidance
₹499₹839
Popular
Digital Product
4.7

Free Resume Template

Download this and edit using Google Docs
FREE
Digital Product
5

Free Referral Template

Use the first one, short and effective.
FREE
Video meeting . 45 mins

College Students Personalized Roadmap

Use your college time wisely
₹399₹999
Video meeting . 15 mins

Quick video call about anything!

Tech, hackathons, networking, relationship...
₹298₹799

About me

Thanks for checking out my profile, I hope you'll find something useful. Actively working towards expanding my knowledge as an Big Data Engineer and working in the field of Big Data. Eagerly demystifying the software development black box. Delivered 10+ sessions for Azure Developer Community and 2 boot camps on AI-900 and 6+ conferences. Hackathons: Atlassian Codeigest 2024 - Winner - ($2,500) Citi Bank Innovation Challenge - 3rd Prize (~$2000) Deepgram X DEV Hackathon 2022 - Winner ($1500) Microsoft Azure Trail Hackathon 2022 - Winner ($1500) Azure Developer Stories 2021 - Winner (Microsoft Surface Go 2) Azure Community Conference 2021 Hackathon - Winner (Microsoft Surface Pro X) Microsoft HackNight 1.0 2019 (Bangalore) - Finalist Aerothon 3.0 2021 by AirBus (Bangalore) - Finalist CODIEcon 2019 by Coviam Technologies (Bangalore) Actively participating in community events of GDG, GCD, TensorFlow User Group and using my leisure to help budding developers by writing technical blogs on Medium and developing open-source projects.

Frequently asked questions

How to learn data engineering from scratch?

Start with SQL and Python, then move to one distributed processing tool (Apache Spark is the industry standard) and one cloud platform. Learn by building — small end-to-end pipelines teach you far more than passive videos. For a structured starting point, Sandy's "Databricks – Big Book of Data Engineering" compiles the essentials in one place, and his 1:1 mentorship helps you sequence the learning around your studies or current job.

What is a realistic data engineering roadmap for college students?

A practical order is: programming fundamentals (Python) and SQL first, then databases and data structures, then one big data tool like Spark plus one cloud such as Azure, and finally 2–3 portfolio projects with a certification. Instead of following a generic timeline, adjust this to your branch, semester, and target companies — Sandy offers a College Students Personalized Roadmap call on Topmate built exactly for this.

Data engineering vs data science — which one should I choose as a fresher?

Data engineering is about building and maintaining the pipelines and platforms that collect, store, and move data, while data science is about analyzing that data and building models on top of it. If you enjoy coding, systems, and SQL-heavy problem solving, lean data engineering; if you prefer statistics, experiments, and storytelling with data, lean data science. This is one of the most common confusions Sandy handles in his quick 1:1 calls, where he matches your strengths to a practical direction instead of a one-size-fits-all answer.

What is a data engineering role, and how is it different from software development?

A data engineering role focuses on designing, building, and optimizing data pipelines — ingesting raw data, transforming it with tools like Spark, Hive, and Airflow, and making it reliable and usable for analysts, data scientists, and applications. Software development builds user-facing or system-level products, while data engineering makes the underlying data fast, clean, and available at scale. Sandy works as a Data Engineer III at Walmart, so his sessions give you a realistic picture of the day-to-day job inside large companies.

Are paid data engineering courses worth it, or can I learn for free?

You can learn data engineering for free using documentation, YouTube, and community blogs, but structured material saves months by giving you an ordered path and practice problems. A sensible middle path is to combine free resources with one affordable structured resource — such as Sandy's Databricks Big Book of Data Engineering or his Master Python notes — before committing to an expensive full-length course.

What kind of data engineering projects should I build to get shortlisted?

Recruiters value end-to-end projects over certificates: take a public dataset, ingest it, build batch pipelines with Python and Spark, orchestrate them with Airflow, and serve the output through a warehouse or dashboard. Add one streaming or incremental-load project if possible, and keep everything on GitHub with clear READMEs. Sandy's 13 hackathon wins — including Atlassian Codeigest 2024 and the Citi Bank Innovation Challenge — show how competitive, real-world builds also make a fresher's resume stand out.

What are the most common data engineering interview questions?

They usually fall into four clusters: SQL (joins, window functions, tuning), coding in Python or Scala, Spark internals (transformations vs actions, shuffles, partitioning, caching), and data modeling or scenario-based pipeline design. Interviewers also dig deep into your projects, so be ready to explain the trade-offs you made. Sandy's resume review sessions focus on getting your projects and keywords past the shortlisting stage before these questions even reach you.

How do freshers get data engineering jobs in India?

The working formula is strong SQL and Python, 2–3 solid pipeline projects, one cloud certification, and referrals — a referral dramatically increases the chance your resume is actually seen. Sandy shares a Free Referral Template and accepts referral requests through priority DMs, which many freshers use after fixing their resume and roadmap in a 1:1 call.

How to learn Apache Spark if I already know Python?

Begin with PySpark, since your Python skills transfer directly: learn the DataFrame API first, then go one level deeper into RDDs, lazy evaluation, shuffles, and partitions. Practice on Databricks Community Edition with realistic datasets rather than toy examples. Sandy's Databricks – Big Book of Data Engineering follows this same progression, and a mentorship call helps when you get stuck on internals.

What is Apache Spark used for in the industry?

Spark is used wherever data becomes too big for a single machine: ETL and batch processing, streaming ingestion, powering analytics and recommendation systems, and preparing data at scale for machine learning. Retail, banking, and tech companies run their daily critical pipelines on it, which is why it appears in almost every data engineer job description alongside Hive and Airflow.

Are Apache Spark and PySpark the same?

No — Apache Spark is the distributed computing engine itself, while PySpark is simply the Python API for using that engine. PySpark code runs on the same Spark core as Scala or Java programs, so the choice of language doesn't change what Spark can do. Since Python is the easiest entry point for most learners, beginners typically start with PySpark and learn Spark internals alongside it.

Which Apache Spark interview questions should I prepare for data engineer roles?

Prioritize transformations vs actions, lazy evaluation, shuffle and partitioning behavior, caching and persistence, join strategies, skewed data handling, and batch vs streaming differences. Many interviewers also give a live PySpark scenario — such as deduplicating or aggregating a huge dataset — so practice writing real code, not just definitions. Sandy works with Spark and Hive at Walmart, which makes his 1:1 sessions useful for realistic, interview-style practice on these topics.

Which Azure certifications for data engineers are worth doing first?

A sensible sequence is one fundamentals cert (AZ-900, or AI-900 if you lean toward AI/ML), followed by the Azure data engineer track, and then a specialty aligned with your target role. Sandy is 6x Azure certified and a Microsoft Certified Trainer who has conducted AI-900 bootcamps, so his mentorship calls are often used to plan the right exam order for a data engineering career.

What is the Azure certification cost in India?

Azure exam fees in India are priced locally and vary by level — fundamentals exams are the cheapest, while role-based exams on the data engineer track cost a few thousand rupees plus GST. Microsoft also regularly provides discount vouchers through student programs, virtual training events, and community challenges, so check the current fee on the official certification page before scheduling. Plan one exam at a time instead of booking several together.

How to renew an Azure certification before it expires?

Azure role-based certifications are valid for one year and can be renewed free of cost by passing an online renewal assessment on Microsoft Learn during the final six months of validity. You receive email reminders, but if a certification lapses you have to take the full exam again, so set your own reminder a couple of months early. With six Azure certifications of his own, Sandy often gets asked in sessions how working professionals manage renewals alongside a full-time job.