Lead Site Reliability Engineer
Teknik, data och digitalt · IT-infrastruktur och säkerhet · Site reliability engineering · DevOps
I korthet
Lead Site Reliability Engineer (SRE) needed to oversee support operations and enhance system performance, availability, and resiliency. Responsibilities include managing a team, monitoring systems, developing incident response procedures, and leading automation efforts.
Ansvarsområden
- Manage a team of support engineers and SREs to provide technical support and address system issues promptly.
- Monitor system performance and reliability metrics, identifying areas for improvement and implementing solutions.
- Collaborate with cross-functional teams to optimize application performance and enhance system reliability.
- Develop and maintain incident response procedures and protocols to minimize system downtime.
- Conduct regular audits and assessments to ensure compliance with industry standards and best practices.
- Lead the implementation of automation tools and processes to streamline support operations and enhance efficiency.
- Provide technical expertise and guidance to team members, promoting a culture of continuous learning and development.
- Plug-in findings into CI/CD pipeline to stop code deployments beyond Beta1.
- Document observations via feedback loop and enable the process to improvise the agent.
- See automated health checks with the built pipeline with the customer's session details.
Krav
- Proficiency in site reliability engineering (SRE) principles and practices.
- Strong background in system administration, networking, and cloud computing.
- Experience with monitoring tools such as Prometheus, Grafana, and ELK stack.
- Knowledge of containerization technologies like Docker and Kubernetes.
- Ability to troubleshoot complex technical issues and perform root cause analysis.
- Excellent communication skills and ability to work collaboratively in a team environment.
- Strong project management and leadership skills to drive initiatives and deliver results efficiently.
- Splunk, ELK, Dynatrace, AppDynamics, Grafana
Önskade kvalifikationer
- Certifications in relevant areas such as AWS Certified DevOps Engineer or Google Professional Cloud DevOps Engineer are a plus.
Förmåner
- Supercharge your potential.
- Find your career.
- Find your spark.
- A place that knows that helping its customers stay on top starts by putting its people first.
#SRE#Site Reliability Engineering#DevOps#Cloud#Support