Senior Site Reliability Engineer Lead

Teknik, data och digitalt · IT-infrastruktur och säkerhet · Site reliability engineering · DevOps

I korthet

We are seeking a Senior Site Reliability Engineer Lead with extensive experience in driving operational excellence, reliability engineering, and observability strategies. This role involves leading major incidents, improving service reliability, and implementing enterprise-wide monitoring and automation frameworks, with a focus on Mainframe and enterprise platform support.

Ansvarsområden

  • Provide Service Reliability Leadership and lead Operational Excellence Programs.
  • Develop Reliability Engineering Strategy and drive Service Maturity Improvement.
  • Lead Major Incident Management and Executive Incident Communications.
  • Oversee Root Cause Elimination and Problem Management Governance.
  • Define and implement Enterprise Monitoring Strategy and Dashboard Governance.
  • Develop and execute Automation and Toil Reduction Programs.
  • Lead Mainframe Operations and z/OS Platform Support.
  • Conduct Operational Readiness Reviews and manage Release Governance.
  • Engage with Engineering, Product, Infrastructure, and Business Stakeholders.
  • Mentor teams, enable Subject Matter Experts, and drive Knowledge Management Programs.

Krav

  • Extensive experience driving operational excellence, reliability engineering, service maturity initiatives, observability strategies, automation programs, and large-scale production support transformations.
  • Proven ability to lead major incidents and influence engineering and business stakeholders.
  • Experience improving service reliability and implementing enterprise-wide monitoring, automation, and operational governance frameworks.
  • Expertise in Splunk, Netcool, xMatters, Domo, Remedy, Jira, Bitbucket, XL Release, Endevor, SDSF, File-AID, Abend-Aid, IDCAMS, DFSORT/SyncSort, JCL, REXX, CICS, and z/OS.
  • Experience with Mainframe & Enterprise Platform Support, including IBM Mainframe Operations and z/OS Platform Support.
  • Knowledge of Release & Operational Readiness Governance, including Operational Readiness Reviews and Production Acceptance Criteria.
  • Strong stakeholder management and cross-functional leadership skills.
  • Experience in Knowledge Management and Team Mentoring.

Önskade kvalifikationer

  • Relevant certifications in Site Reliability Engineering (SRE) or Cloud Services are a plus.

Förmåner

  • Supercharge your potential at HCLTech.
  • Find your career and your spark at a company that puts people first.
  • Opportunity to work with a global technology leader with consolidated revenues of $14.8 billion.
#Site Reliability Engineering#SRE#Operational Excellence#Reliability Engineering#Service Maturity#Observability#Automation#Production Support#Incident Management#Mainframe#z/OS#CICS#JCL#REXX
HCLTech Logo

Företag

HCLTech

Publicerade jobb

för 2 veckor sedan

Anställningstyp

Heltid

Arbetsform

På plats

Erfarenhetsnivå

Senior

Platser

Pune, India

Sökande

Ansök tidigt