Sr Technical Lead (Support & Operations)

Teknik, data och digitalt · IT-infrastruktur och säkerhet · DevOps · Site reliability engineering · Systemteknik

I korthet

We are seeking an experienced Prometheus Lead (L3) to oversee the architecture, administration, and optimization of the enterprise Prometheus monitoring platform. This role involves designing metrics collection, monitoring strategies, and alerting mechanisms across diverse environments, requiring strong expertise in Prometheus and related technologies.

Ansvarsområden

  • Lead the administration, governance, and lifecycle management of the Prometheus monitoring platform.
  • Design and maintain enterprise-wide monitoring solutions using Prometheus.
  • Define monitoring standards, metric collection policies, alerting strategies, and best practices.
  • Implement scalable monitoring architectures across infrastructure, applications, cloud, databases, and container platforms.
  • Configure and manage service discovery mechanisms and exporters.
  • Develop monitoring frameworks to provide visibility into system health, performance, availability, and capacity.
  • Establish monitoring baselines, thresholds, and alerting standards.
  • Collaborate with infrastructure, cloud, application, and SRE teams to onboard new services into monitoring.
  • Drive platform optimization, scalability improvements, and monitoring standardization.
  • Provide technical leadership and mentorship to monitoring engineers and platform administrators.
  • Create technical documentation, operational procedures, and knowledge articles.

Krav

  • Strong hands-on experience with Prometheus administration and monitoring architecture.
  • Deep understanding of metrics collection, monitoring frameworks, and alerting concepts.
  • Experience with PromQL query development and optimization.
  • Strong knowledge of Prometheus federation and scalability concepts.
  • Experience configuring exporters and monitoring integrations.
  • Knowledge of Linux administration and troubleshooting.
  • Understanding of infrastructure monitoring, application monitoring, and cloud monitoring.
  • Experience with containerized environments and Kubernetes monitoring.
  • Scripting knowledge in Python, Bash, PowerShell, YAML.
  • Excellent troubleshooting, analytical, and problem-solving skills.
  • Strong communication and stakeholder management capabilities.
  • Experience leading technical teams and enterprise monitoring initiatives.
  • 9+ years of experience in enterprise monitoring, metrics collection, infrastructure monitoring, performance monitoring, alerting, and Prometheus platform administration.

Önskade kvalifikationer

  • Experience with Kubernetes and OpenShift environments.
  • Knowledge of Grafana dashboard integration.
  • Exposure to cloud monitoring technologies (Azure, AWS, GCP).
  • Experience with ServiceNow integration and ITSM processes.
  • Knowledge of SRE and observability practices.
  • Exposure to OpenTelemetry and modern monitoring architectures.
  • Experience supporting large-scale enterprise monitoring environments.
  • Understanding of capacity planning and performance engineering.
  • Bachelor’s degree in engineering, Computer Science, Information Technology, or related discipline.
  • Prometheus Certified Professional (if applicable)
  • Kubernetes Certifications (CKA/CKAD)
  • Linux Administration Certifications
  • Cloud Certifications (Azure/AWS/GCP)
  • ITIL Foundation

Förmåner

  • You'll supercharge your potential.
  • You'll find your career.
  • And you'll find your spark.
#Prometheus#Support#Operations#Monitoring#Alerting#SRE#Observability#Kubernetes#Cloud
HCLTech Logo

Företag

HCLTech

Publicerade jobb

för 4 veckor sedan

Anställningstyp

Heltid

Arbetsform

På plats

Erfarenhetsnivå

Senior

Platser

Greater Noida, India

Kvalifikation

Kandidatexamen

Sökande

Ansök tidigt