Senior Staff Infrastructure & Site Reliability Engineer – Datacentre AI Engineering
Teknik, data och digitalt · Data, AI och analys · Data engineering · Maskininlärning · Site reliability engineering
I korthet
Qualcomm is hiring a Senior Staff Infrastructure & Site Reliability Engineer for their Datacentre AI Engineering team in Riyadh, Saudi Arabia. This role focuses on designing, operating, and improving large-scale AI inference systems, ensuring reliability, scalability, and production readiness for advanced machine-learning workloads. The ideal candidate will have 12+ years of experience in SRE with strong systems and software engineering fundamentals.
Ansvarsområden
- Design, deploy, and operate large-scale AI inference systems supporting critical AI workloads.
- Ensure reliability, availability, and scalability of Qualcomm datacenter AI clusters.
- Develop and maintain software tools and support infrastructure around AI software stacks.
- Build, deploy, and operate components supporting LLM inference, agentic AI workflows, and AI services.
- Work with models, systems, and software teams to improve model performance on AI100 deployments.
- Identify and implement optimizations for workloads running on multi-SoC and multi-card systems.
- Apply SRE fundamentals including monitoring, alerting, incident response, and performance optimization.
- Support production ML systems using MLOps tools and operational best practices.
- Contribute to incident reviews, operational documentation, and continuous reliability improvements.
- Build and maintain observability tools, dashboards, and alerts to monitor system health and reliability.
- Monitor infrastructure and services using tools such as Prometheus, Grafana, CloudWatch, and custom telemetry.
- Create and maintain technical documentation, runbooks, and knowledge-base articles.
- Develop automation to reduce manual operational tasks and improve system reliability.
- Support CI/CD pipelines for AI service and agent deployment.
- Apply Infrastructure-as-Code practices using tools such as Terraform and Ansible.
Krav
- Experience working with AI/ML workloads such as LLMs, NLP, Vision, Audio, or Recommendation systems.
- Understand ML inference concepts including batching, token streaming, and performance considerations.
- Hands-on experience with PyTorch and familiarity with modern ML frameworks.
- Familiarity with distributed inference, checkpointing, and accelerator-based compute environments.
- Experience supporting AI or ML applications in production environments.
- Familiarity with LLM inference pipelines and AI service operations.
- Strong programming skills in Python with experience building and supporting production systems.
- Experience with scripting and automation using Python and Bash.
- Familiarity with configuration management and orchestration tools.
- Strong Linux fundamentals include shell, containers, system services, and networking basics (DNS, TLS, HTTP/gRPC).
- Experience working with cluster schedulers such as Slurm or equivalent systems.
- Experience operating distributed systems with high availability and fault tolerance.
- Hands-on experience with monitoring and logging tools such as Prometheus, Grafana, ELK, or Loki.
- Understanding of incident management, service health metrics, and system reliability monitoring.
- Solid understanding of SDLC, release processes, and operational reliability practices.
- Familiarity with CI/CD pipelines and Infrastructure-as-Code tools.
Önskade kvalifikationer
- Experience with GenAI, Agentic AI systems, or LLM orchestration frameworks.
- Exposure to LangChain, AutoGen, or RAG-based systems.
- Experience with additional ML frameworks such as TensorFlow, JAX, or Ray.
- Knowledge of GPU/accelerator-based systems and high-performance networking (RDMA, InfiniBand, RoCE).
- Experience with advanced MLOps workflows or large-scale AI platform operations.
Förmåner
- Salary including housing & transport allowance
- Stock (RSU's) and performance related bonus
- 16 weeks fully paid Maternity Leave
- 6 weeks fully paid Paternity Leave
- Employee stock purchase scheme
- Child Education Allowance
- Relocation and immigration support (if needed)
- Life and Medical Insurance
- Live+ Well Reimbursement for health and recreational membership fees
#AI#Data Centre#SRE#Infrastructure#Machine Learning#Riyadh#Saudi Arabia#Vision 2030#LLM#MLOps#DevOps#Python#Linux#Cloud#Networking