Senior Solution Architect
Technology, Data & Digital · IT Infrastructure & Security · Site Reliability Engineering · DevOps · Cloud Engineering
In short
We are seeking an experienced SRE leader to drive enterprise-wide reliability transformation across various platforms. This role requires technical expertise and strategic leadership in modernizing operations through SRE, AI-driven observability, and automation practices. You will partner with multiple teams to institutionalize modern Site Reliability Engineering practices at scale.
Responsibilities
- Lead and drive enterprise-wide SRE transformation initiatives.
- Assess operational maturity and define SRE adoption roadmaps and strategies.
- Coach and mentor teams on SRE principles, SLI/SLO frameworks, and reliability engineering.
- Establish and institutionalize reliability metrics and service health management practices.
- Define and govern enterprise observability strategies leveraging Splunk.
- Lead implementation of enterprise observability capabilities.
- Guide teams in implementing Incident Management, Problem Management, and RCA.
- Drive continuous improvement initiatives to reduce MTTD and MTTR.
- Promote automation-first operations through toil reduction and intelligent automation.
- Advise and guide teams on implementing AIOps capabilities.
- Collaborate with stakeholders to embed reliability into technology delivery and operational processes.
- Drive adoption of Agile and Scrum practices.
- Measure and report SRE adoption and reliability KPIs to leadership.
Requirements
- Deep expertise in Site Reliability Engineering (SRE) including SLI, SLO, Error Budgets, reliability governance, and service health management with 13+ years of experience.
- Proven experience leading enterprise SRE transformation programs and coaching engineering and operations teams.
- Strong hands-on expertise with Splunk Observability Cloud, Splunk ITSI, and Splunk Enterprise.
- Strong understanding of enterprise observability architecture, telemetry strategy, monitoring standards, and operational analytics.
- Experience driving Incident Management, Root Cause Analysis (RCA), and Blameless Postmortems.
- Strong track record in Automation and Toil Reduction initiatives (Ansible, Python, RPA and other automation platforms).
- Experience implementing AIOps and Intelligent Operations practices.
- Hands-on experience with cloud platforms such as Azure, AWS, and GCP.
- Experience institutionalizing SRE practices for various business applications including custom applications, SAP, Salesforce, SaaS/COTS, and business-critical enterprise applications.
- Drive continuous improvement through Agile and Scrum practices.
- Strong stakeholder management, leadership, communication, and mentoring skills.
Desired Qualifications
- Knowledge of DevOps practices, CI/CD pipelines, GitOps, and release automation.
- Experience with Reliability Engineering practices including Resilience Testing and Chaos Engineering.
- Experience with Agentic AI, AI Agents, and AI-driven solutions for SRE, DevOps, Automation, and AMS operations.
- Experience with Platform Engineering, Internal Developer Platforms (IDP), and self-service engineering models.
- SRE, Splunk, Cloud, Observability, or related industry certifications.
Benefits
- Supercharge your potential at HCLTech.
- Find your career and spark at HCLTech.
- HCLTech puts its people first.
- Global technology company with over 223,000 people across 60 countries.
- Delivering industry-leading capabilities centered around digital, engineering, cloud, and AI.
- Work with clients across all major verticals.
- Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.
#SRE#Site Reliability Engineering#Observability#Cloud#AI#Automation#Splunk