Lead Site Reliability Engineer
Technology, Data & Digital · IT Infrastructure & Security · Site Reliability Engineering · DevOps
In short
Lead Site Reliability Engineer (SRE) needed to oversee support operations and enhance system performance, availability, and resiliency. Responsibilities include managing a team, monitoring systems, developing incident response procedures, and leading automation efforts.
Responsibilities
- Manage a team of support engineers and SREs to provide technical support and address system issues promptly.
- Monitor system performance and reliability metrics, identifying areas for improvement and implementing solutions.
- Collaborate with cross-functional teams to optimize application performance and enhance system reliability.
- Develop and maintain incident response procedures and protocols to minimize system downtime.
- Conduct regular audits and assessments to ensure compliance with industry standards and best practices.
- Lead the implementation of automation tools and processes to streamline support operations and enhance efficiency.
- Provide technical expertise and guidance to team members, promoting a culture of continuous learning and development.
- Plug-in findings into CI/CD pipeline to stop code deployments beyond Beta1.
- Document observations via feedback loop and enable the process to improvise the agent.
- See automated health checks with the built pipeline with the customer's session details.
Requirements
- Proficiency in site reliability engineering (SRE) principles and practices.
- Strong background in system administration, networking, and cloud computing.
- Experience with monitoring tools such as Prometheus, Grafana, and ELK stack.
- Knowledge of containerization technologies like Docker and Kubernetes.
- Ability to troubleshoot complex technical issues and perform root cause analysis.
- Excellent communication skills and ability to work collaboratively in a team environment.
- Strong project management and leadership skills to drive initiatives and deliver results efficiently.
- Splunk, ELK, Dynatrace, AppDynamics, Grafana
Desired Qualifications
- Certifications in relevant areas such as AWS Certified DevOps Engineer or Google Professional Cloud DevOps Engineer are a plus.
Benefits
- Supercharge your potential.
- Find your career.
- Find your spark.
- A place that knows that helping its customers stay on top starts by putting its people first.
#SRE#Site Reliability Engineering#DevOps#Cloud#Support