CMPE 189: site reliability engineering for cloud
fall 2026 — MW 1:30 — taught by ben reed
ever wonder how google, netflix, and friends keep their services running while deploying new code thousands of times a day? that’s site reliability engineering, and it’s one of the most in-demand skills in industry right now.
this course was developed by google, and this is the first time it is being taught anywhere. you will be part of the very first class to take it.
this course is a hands-on introduction to SRE. you will learn:
- SLIs, SLOs, and error budgets — measuring reliability and balancing reliability with progress
- datacenter infrastructure — what the cloud is actually made of
- virtualization and storage — carving up machines and keeping data safe
- scalability — making systems that survive success
- application lifecycle management — from build to deployment to observability
- capacity management — planning for growth without burning money
- cross-application security — keeping the whole system safe, not just one app
the projects are hands-on and use industry best practices, RPC frameworks, and the observability tools used by real SRE teams. you will leave this course knowing how production systems are built and kept alive.
see why you should register
how to register
note: there are a couple of different CMPE 189 courses this fall. look for the one taught by ben reed MW 1:30 for SRE. (the other one looks cool too ;)