Running systems in production means being on top of reliability, observability, and incidents — but SRE tooling can feel complex. In this 45-minute 1:1 session, I'll help you understand the SRE mindset and build a practical observability setup for your systems.
What we'll cover:
SRE fundamentals: SLIs, SLOs, error budgets
Monitoring with Prometheus: metrics, exporters, alerting rules
Dashboards in Grafana: visualising system health
Logging and tracing overview (ELK, Loki, Jaeger)
Incident response basics and on-call best practices
Ideal for: DevOps engineers, platform engineers, and developers who want to move beyond basic monitoring and build truly observable, reliable production systems.