Lead Site Reliability Engineer

Teknik, data och digitalt · IT-infrastruktur och säkerhet · Site reliability engineering · DevOps

I korthet

Lead Site Reliability Engineer (SRE) needed to oversee support operations and enhance system performance, availability, and resiliency. Responsibilities include managing a team, monitoring systems, developing incident response procedures, and leading automation efforts.

Ansvarsområden

  • Manage a team of support engineers and SREs to provide technical support and address system issues promptly.
  • Monitor system performance and reliability metrics, identifying areas for improvement and implementing solutions.
  • Collaborate with cross-functional teams to optimize application performance and enhance system reliability.
  • Develop and maintain incident response procedures and protocols to minimize system downtime.
  • Conduct regular audits and assessments to ensure compliance with industry standards and best practices.
  • Lead the implementation of automation tools and processes to streamline support operations and enhance efficiency.
  • Provide technical expertise and guidance to team members, promoting a culture of continuous learning and development.
  • Plug-in findings into CI/CD pipeline to stop code deployments beyond Beta1.
  • Document observations via feedback loop and enable the process to improvise the agent.
  • See automated health checks with the built pipeline with the customer's session details.

Krav

  • Proficiency in site reliability engineering (SRE) principles and practices.
  • Strong background in system administration, networking, and cloud computing.
  • Experience with monitoring tools such as Prometheus, Grafana, and ELK stack.
  • Knowledge of containerization technologies like Docker and Kubernetes.
  • Ability to troubleshoot complex technical issues and perform root cause analysis.
  • Excellent communication skills and ability to work collaboratively in a team environment.
  • Strong project management and leadership skills to drive initiatives and deliver results efficiently.
  • Splunk, ELK, Dynatrace, AppDynamics, Grafana

Önskade kvalifikationer

  • Certifications in relevant areas such as AWS Certified DevOps Engineer or Google Professional Cloud DevOps Engineer are a plus.

Förmåner

  • Supercharge your potential.
  • Find your career.
  • Find your spark.
  • A place that knows that helping its customers stay on top starts by putting its people first.
#SRE#Site Reliability Engineering#DevOps#Cloud#Support
HCLTech Logo

Företag

HCLTech

Publicerade jobb

för 2 veckor sedan

Anställningstyp

Heltid

Arbetsform

På plats

Erfarenhetsnivå

Senior

Platser

Hyderabad, India

Sökande

Ansök tidigt