Site Reliability Manager
Teknik, data och digitalt · IT-infrastruktur och säkerhet · Site reliability engineering · Mjukvaruutveckling · Systemteknik
I korthet
Google is seeking a Site Reliability Manager in Bengaluru, India to lead a team of 6-10 SREs supporting enterprise services. This role involves developing roadmaps, ensuring service reliability, and scaling systems through automation. The ideal candidate has a Bachelor's degree in CS or equivalent, 5 years of experience with distributed systems and Kubernetes, and 5 years of people management experience.
Ansvarsområden
- Manage a team of 6-10 site reliability engineers supporting Google’s enterprise services.
- Develop roadmaps, planning, objectives and key results (OKRs) to move forward the maturity of the managed services.
- Engage in and improve the whole lifecycle of services from inception and design, through deployment, operation and refinement.
- Support services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity planning and launch reviews.
- Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
- Scale systems sustainably through mechanisms like automation, and evolve systems by pushing for changes that improve reliability and velocity.
- Practice sustainable incident response ensuring services meet their service level objectives.
Krav
- Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience.
- 5 years of experience building or managing distributed systems or cloud infrastructure, with a focus on Kubernetes.
- 5 years of experience in people management.
- Experience with site reliability engineering, system design, distributed computing.
Önskade kvalifikationer
- 5 years of experience in people management, with managing distributed, multi-site teams through engineering managers or tech leads.
- Experience in Enterprise tooling and technology.
- Experience in Systems, Applications, and Products (SAP) or other Enterprise Resource Planning (ERP) systems.
Förmåner
- Opportunity to manage complex challenges of scale unique to Google Cloud.
- Use expertise in coding, algorithms, complexity analysis and large-scale system design.
- Collaborate in a culture of intellectual curiosity, problem solving and openness.
- Work in a blame-free environment that promotes self-direction.
- Supportive and mentorship-driven environment for learning and growth.
#Site Reliability Engineering#SRE#Kubernetes#Cloud Infrastructure#Distributed Systems#People Management#System Design#Enterprise Tooling#SAP#ERP