Senior DevOps Engineer – Observability Platform
Teknik, data och digitalt · IT-infrastruktur och säkerhet · DevOps · Mjukvaruutveckling
I korthet
As a Senior DevOps Engineer specializing in Observability Platforms, you will build and maintain scalable, reliable infrastructure and deployment pipelines on Kubernetes and AWS. This role emphasizes metrics, logs, and traces, collaborating with development teams to enhance velocity while ensuring reliability, security, and performance. You will design and operate a self-service observability platform, instrument applications, define SLOs, and optimize system performance.
Ansvarsområden
- Build and maintain scalable, reliable infrastructure and deployment pipelines with an emphasis on observability (metrics, logs, traces) across systems running on Kubernetes and AWS.
- Work closely with development teams to improve development velocity while ensuring system reliability, security, and performance.
- Provide a standardized observability platform for internal engineering teams and customer-facing services, offering deep visibility into system health, performance, and reliability.
- Design, implement, and maintain cloud-based infrastructure using Infrastructure as Code principles.
- Develop automation scripts and tools to streamline operations and eliminate manual processes.
- Manage containerization strategies and orchestration using Docker and Kubernetes.
- Design, build, and operate a standardized, self-service metrics, logs, and tracing platform (Prometheus, Grafana, Loki, OpenTelemetry) serving both internal teams and external, customer-facing services running on Kubernetes and AWS.
- Partner with engineering teams to instrument applications and infrastructure, standardizing telemetry collection with OpenTelemetry.
- Define and maintain SLIs/SLOs and error budgets, build actionable dashboards, and tune alerting to maximize signal and reduce noise.
- Use observability data to analyze and optimize system performance, scalability, and cost-efficiency.
- Create and maintain thorough documentation for infrastructure, deployment processes, and operational procedures.
- Provide second-tier engineering escalation during business hours and own the telemetry, SLO, and alerting tooling that powers incident detection and reduces MTTD/MTTR.
Krav
- Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
- OR Master's degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience.
- OR PhD in Engineering, Information Systems, Computer Science, or related field.
- 2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.
- 5+ years of experience in DevOps, Observability, or similar roles, including hands-on production experience operating Kubernetes based stack.
- Extensive hands-on experience with AWS, including its observability services (CloudWatch, X-Ray, Amazon Managed Service for Prometheus, Amazon Managed Grafana).
- Proficiency with Terraform, AWS CloudFormation, or similar IaC tools.
- Advanced knowledge of Docker and Kubernetes ecosystem.
- Hands-on experience with Prometheus, Grafana, Loki, Tempo or Jaeger, OpenTelemetry, and Alertmanager; experience scaling metrics storage with Thanos, Mimir, or Cortex.
- Strong coding skills in Python, Bash, or Go.
- Excellent analytical and troubleshooting skills.
- Strong verbal and written communication skills.
- Ability to work effectively in a team environment and collaborate with cross-functional teams.
- Proven leadership skills and the ability to mentor junior engineers.
- Comfortable working in a fully distributed, offshore setup and collaborating effectively with development teams across multiple locations.
Önskade kvalifikationer
- Strong background in software development with security focus
Förmåner
- Qualcomm is an equal opportunity employer. If you are an individual with a disability and need an accommodation during the application/hiring process, rest assured that Qualcomm is committed to providing an accessible process.
- Upon request, Qualcomm will provide reasonable accommodations to support individuals with disabilities to be able participate in the hiring process.
- Qualcomm is also committed to making our workplace accessible for individuals with disabilities.
Färdigheter
KubernetesAWSPrometheusGrafanaLokiOpenTelemetryTerraformPythonBashGoDockerCloudWatchX-RayAmazon Managed Service for PrometheusAmazon Managed GrafanaAWS CloudFormationTempoJaegerAlertmanagerThanosMimirCortex
#DevOps#Observability#Kubernetes#AWS#Cloud#Infrastructure#Automation#Containerization#Metrics#Logs#Tracing#Prometheus#Grafana#Loki#OpenTelemetry#Terraform#Python#Bash#Go
Enabling a world where everyone and everything can be intelligently connected.\n\n
Företag
QualcommPublicerade jobb
för 20 timmar sedan
Anställningstyp
Heltid
Arbetsform
Distans
Erfarenhetsnivå
Senior
Platser
Chennai, India
Kvalifikation
Kandidatexamen, Masterexamen, Doktorand
Sökande
Ansök tidigt
Liknande jobb
Amazon
Stockholm, SwedenPrincipal Account Manager, Telecoms
Heltid Ansök tidigt Publicerad för 16 timmar sedan
Ericsson
Stockholm, SwedenVerification Engineer
Heltid Ansök tidigt Publicerad för 15 timmar sedan
IKEA
Malmö, SwedenSoftware Engineer | Range Operations
Heltid Ansök tidigt Publicerad för 13 timmar sedan
Ericsson
Lund, SwedenMaster Thesis: LLM-based log processing and Anomaly Detection in Cloud Orchestartion Systems
Praktik Ansök tidigt Publicerad för 10 timmar sedan
Google
London, United Kingdom +1 flerCustomer Engineer, Applied AI
1,36–1,39 mn kr/år Heltid Ansök tidigt Publicerad för 8 timmar sedan
Microsoft
United KingdomSecurity Research IC5
93,5–162 tn GBP/år Heltid Ansök tidigt Publicerad för 4 dagar sedan
