Senior Staff Infrastructure & Site Reliability Engineer – Datacentre AI Engineering

Technology, Data & Digital · Data, AI & Analytics · Data Engineering · Machine Learning · Site Reliability Engineering

In short

Qualcomm is hiring a Senior Staff Infrastructure & Site Reliability Engineer for their Datacentre AI Engineering team in Riyadh, Saudi Arabia. This role focuses on designing, operating, and improving large-scale AI inference systems, ensuring reliability, scalability, and production readiness for advanced machine-learning workloads. The ideal candidate will have 12+ years of experience in SRE with strong systems and software engineering fundamentals.

Responsibilities

  • Design, deploy, and operate large-scale AI inference systems supporting critical AI workloads.
  • Ensure reliability, availability, and scalability of Qualcomm datacenter AI clusters.
  • Develop and maintain software tools and support infrastructure around AI software stacks.
  • Build, deploy, and operate components supporting LLM inference, agentic AI workflows, and AI services.
  • Work with models, systems, and software teams to improve model performance on AI100 deployments.
  • Identify and implement optimizations for workloads running on multi-SoC and multi-card systems.
  • Apply SRE fundamentals including monitoring, alerting, incident response, and performance optimization.
  • Support production ML systems using MLOps tools and operational best practices.
  • Contribute to incident reviews, operational documentation, and continuous reliability improvements.
  • Build and maintain observability tools, dashboards, and alerts to monitor system health and reliability.
  • Monitor infrastructure and services using tools such as Prometheus, Grafana, CloudWatch, and custom telemetry.
  • Create and maintain technical documentation, runbooks, and knowledge-base articles.
  • Develop automation to reduce manual operational tasks and improve system reliability.
  • Support CI/CD pipelines for AI service and agent deployment.
  • Apply Infrastructure-as-Code practices using tools such as Terraform and Ansible.

Requirements

  • Experience working with AI/ML workloads such as LLMs, NLP, Vision, Audio, or Recommendation systems.
  • Understand ML inference concepts including batching, token streaming, and performance considerations.
  • Hands-on experience with PyTorch and familiarity with modern ML frameworks.
  • Familiarity with distributed inference, checkpointing, and accelerator-based compute environments.
  • Experience supporting AI or ML applications in production environments.
  • Familiarity with LLM inference pipelines and AI service operations.
  • Strong programming skills in Python with experience building and supporting production systems.
  • Experience with scripting and automation using Python and Bash.
  • Familiarity with configuration management and orchestration tools.
  • Strong Linux fundamentals include shell, containers, system services, and networking basics (DNS, TLS, HTTP/gRPC).
  • Experience working with cluster schedulers such as Slurm or equivalent systems.
  • Experience operating distributed systems with high availability and fault tolerance.
  • Hands-on experience with monitoring and logging tools such as Prometheus, Grafana, ELK, or Loki.
  • Understanding of incident management, service health metrics, and system reliability monitoring.
  • Solid understanding of SDLC, release processes, and operational reliability practices.
  • Familiarity with CI/CD pipelines and Infrastructure-as-Code tools.

Desired Qualifications

  • Experience with GenAI, Agentic AI systems, or LLM orchestration frameworks.
  • Exposure to LangChain, AutoGen, or RAG-based systems.
  • Experience with additional ML frameworks such as TensorFlow, JAX, or Ray.
  • Knowledge of GPU/accelerator-based systems and high-performance networking (RDMA, InfiniBand, RoCE).
  • Experience with advanced MLOps workflows or large-scale AI platform operations.

Benefits

  • Salary including housing & transport allowance
  • Stock (RSU's) and performance related bonus
  • 16 weeks fully paid Maternity Leave
  • 6 weeks fully paid Paternity Leave
  • Employee stock purchase scheme
  • Child Education Allowance
  • Relocation and immigration support (if needed)
  • Life and Medical Insurance
  • Live+ Well Reimbursement for health and recreational membership fees
#AI#Data Centre#SRE#Infrastructure#Machine Learning#Riyadh#Saudi Arabia#Vision 2030#LLM#MLOps#DevOps#Python#Linux#Cloud#Networking
Qualcomm Logo

Company

Qualcomm

Job Posted

3 weeks ago

Employment Type

Full Time

WorkMode

On Site

Experience Level

Executive

Locations

Riyadh, Saudi Arabia

Qualification

Bachelor, Master, Doctoral

Applicants

Be an early applicant