Senior Site Reliability Engineer Lead
Teknik, data och digitalt · IT-infrastruktur och säkerhet · Site reliability engineering · DevOps
I korthet
We are seeking a Senior Site Reliability Engineer Lead with extensive experience in driving operational excellence, reliability engineering, and observability strategies. This role involves leading major incidents, improving service reliability, and implementing enterprise-wide monitoring and automation frameworks, with a focus on Mainframe and enterprise platform support.
Ansvarsområden
- Provide Service Reliability Leadership and lead Operational Excellence Programs.
- Develop Reliability Engineering Strategy and drive Service Maturity Improvement.
- Lead Major Incident Management and Executive Incident Communications.
- Oversee Root Cause Elimination and Problem Management Governance.
- Define and implement Enterprise Monitoring Strategy and Dashboard Governance.
- Develop and execute Automation and Toil Reduction Programs.
- Lead Mainframe Operations and z/OS Platform Support.
- Conduct Operational Readiness Reviews and manage Release Governance.
- Engage with Engineering, Product, Infrastructure, and Business Stakeholders.
- Mentor teams, enable Subject Matter Experts, and drive Knowledge Management Programs.
Krav
- Extensive experience driving operational excellence, reliability engineering, service maturity initiatives, observability strategies, automation programs, and large-scale production support transformations.
- Proven ability to lead major incidents and influence engineering and business stakeholders.
- Experience improving service reliability and implementing enterprise-wide monitoring, automation, and operational governance frameworks.
- Expertise in Splunk, Netcool, xMatters, Domo, Remedy, Jira, Bitbucket, XL Release, Endevor, SDSF, File-AID, Abend-Aid, IDCAMS, DFSORT/SyncSort, JCL, REXX, CICS, and z/OS.
- Experience with Mainframe & Enterprise Platform Support, including IBM Mainframe Operations and z/OS Platform Support.
- Knowledge of Release & Operational Readiness Governance, including Operational Readiness Reviews and Production Acceptance Criteria.
- Strong stakeholder management and cross-functional leadership skills.
- Experience in Knowledge Management and Team Mentoring.
Önskade kvalifikationer
- Relevant certifications in Site Reliability Engineering (SRE) or Cloud Services are a plus.
Förmåner
- Supercharge your potential at HCLTech.
- Find your career and your spark at a company that puts people first.
- Opportunity to work with a global technology leader with consolidated revenues of $14.8 billion.
#Site Reliability Engineering#SRE#Operational Excellence#Reliability Engineering#Service Maturity#Observability#Automation#Production Support#Incident Management#Mainframe#z/OS#CICS#JCL#REXX