Senior Technical Specialist
Teknik, data och digitalt · IT-infrastruktur och säkerhet · Site reliability engineering · Molnteknik · Mjukvaruutveckling
I korthet
HCLTech seeks a Senior Technical Specialist in Hyderabad, Telangana, with strong experience in Site Reliability Engineering and Resilience Engineering within business-critical environments. The role involves defining and implementing resilience, recoverability, and performance testing, leveraging cloud-native platforms like GCP and GKE, and utilizing observability tools.
Ansvarsområden
- Define and implement resilience, recoverability, capacity, performance, disaster recovery, failover and operational readiness testing.
- Assess production risk, challenge weak controls and drive clear remediation across engineering and supplier teams.
- Develop and maintain documentation and engage stakeholders regarding resilience evidence and operational readiness.
Krav
- Strong experience in Site Reliability Engineering, Resilience Engineering, Platform Engineering, Production Engineering or Non-Functional Testing within business-critical environments.
- Strong understanding of SRE principles including SLIs, SLOs, error budgets, reliability dashboards, toil reduction and evidence-based service improvement.
- Experience working with cloud-native platforms, preferably GCP and GKE.
- Good understanding of distributed systems, Kafka, databases, APIs, service mesh, networking and microservice reliability patterns.
- Experience with observability tooling, dashboards, alerting, incident analysis, problem management and post-incident improvement.
Önskade kvalifikationer
- Banking or core banking platform experience.
- Experience working with Thought Machine Vault.
- Experience supporting important business services, operational resilience, regulatory evidence or CAT A/B type service classification.
- Experience with chaos testing, game days, performance testing, capacity modelling or disaster recovery exercises.
- Experience with Dynatrace, OpenTelemetry, GCP operations tooling, Kafka, AlloyDB, PostgreSQL or Kubernetes.
- Understanding of change management, service readiness, risk management and control evidence within a regulated environment.
Förmåner
- Supercharge your potential at HCLTech.
- Find your career.
- Find your spark.
- A place that knows that helping its customers stay on top starts by putting its people first.
#Reliability Engineering#Resilience Engineering#Site Reliability Engineering#Platform Engineering#Production Engineering#Non-Functional Testing#Cloud Native#GCP#GKE#Distributed Systems#Kafka#Databases#APIs#Service Mesh#Microservices#Observability#Alerting#Incident Management#Problem Management#Thought Machine Vault#Chaos Testing#Performance Testing#Capacity Planning#Disaster Recovery#Dynatrace#OpenTelemetry#AlloyDB#PostgreSQL#Kubernetes#Change Management#Risk Management