Lead Site Reliability Engineer

Technology, Data & Digital · IT Infrastructure & Security · Site Reliability Engineering · DevOps

In short

Lead Site Reliability Engineer (SRE) needed to oversee support operations and enhance system performance, availability, and resiliency. Responsibilities include managing a team, monitoring systems, developing incident response procedures, and leading automation efforts.

Responsibilities

  • Manage a team of support engineers and SREs to provide technical support and address system issues promptly.
  • Monitor system performance and reliability metrics, identifying areas for improvement and implementing solutions.
  • Collaborate with cross-functional teams to optimize application performance and enhance system reliability.
  • Develop and maintain incident response procedures and protocols to minimize system downtime.
  • Conduct regular audits and assessments to ensure compliance with industry standards and best practices.
  • Lead the implementation of automation tools and processes to streamline support operations and enhance efficiency.
  • Provide technical expertise and guidance to team members, promoting a culture of continuous learning and development.
  • Plug-in findings into CI/CD pipeline to stop code deployments beyond Beta1.
  • Document observations via feedback loop and enable the process to improvise the agent.
  • See automated health checks with the built pipeline with the customer's session details.

Requirements

  • Proficiency in site reliability engineering (SRE) principles and practices.
  • Strong background in system administration, networking, and cloud computing.
  • Experience with monitoring tools such as Prometheus, Grafana, and ELK stack.
  • Knowledge of containerization technologies like Docker and Kubernetes.
  • Ability to troubleshoot complex technical issues and perform root cause analysis.
  • Excellent communication skills and ability to work collaboratively in a team environment.
  • Strong project management and leadership skills to drive initiatives and deliver results efficiently.
  • Splunk, ELK, Dynatrace, AppDynamics, Grafana

Desired Qualifications

  • Certifications in relevant areas such as AWS Certified DevOps Engineer or Google Professional Cloud DevOps Engineer are a plus.

Benefits

  • Supercharge your potential.
  • Find your career.
  • Find your spark.
  • A place that knows that helping its customers stay on top starts by putting its people first.
#SRE#Site Reliability Engineering#DevOps#Cloud#Support
HCLTech Logo

Company

HCLTech

Job Posted

2 weeks ago

Employment Type

Full Time

WorkMode

On Site

Experience Level

Senior

Locations

Hyderabad, India

Applicants

Be an early applicant