Site Reliability Engineer
Own uptime for cloud-native services
Northbridge Cloud Systems · Bengaluru
Job Description
Northbridge Cloud Systems is looking for a Site Reliability Engineer to own uptime and performance for a fleet of customer-facing microservices running at meaningful scale. You will build and maintain infrastructure-as-code using Terraform, manage production Kubernetes clusters end to end, and design alerting and observability pipelines using Prometheus and Grafana so issues surface long before customers notice them. You will participate in a shared on-call rotation, own incident response from detection through resolution, and lead blameless post-incident reviews that actually change how the team builds, not just how it reacts. A big part of the job is proactive: identifying single points of failure before they cause an outage, automating repetitive operational toil, and pushing back on designs that trade reliability for short-term speed. You will work closely with product engineering teams from the design phase onward, so reliability gets baked into new services rather than retrofitted after a painful incident. Strong hands-on Kubernetes and AWS experience is required, along with genuine comfort being paged at odd hours and communicating clearly under pressure.
Designation: SRE
Location: Bengaluru
Industry: Cloud Infrastructure