AI is changing how we test software—and traditional validation methods are no longer enough. In this course, you’ll learn how to evaluate Large Language Models (LLMs) in a structured, practical, and industry-relevant way.
We start from the fundamentals of LLM evaluation, making it beginner-friendly even if you're new to GenAI. Then, we explore key evaluation metrics such as correctness, relevance, coherence, and custom scoring approaches used in real-world AI systems.
You’ll also understand the shift in testing mindset in the AI era—why deterministic testing falls short and how evaluation strategies need to evolve.
A major part of the course focuses on Promptfoo, a powerful tool for LLM evaluation. You’ll learn
Finally, we bring everything together with a live demo + mini project, where you’ll see how to evaluate prompts and LLM outputs in practice.