Tower Lead - Red Hat Enterprise Linux, Red Hat Satellite, Red Hat Cluster
Teknik, data och digitalt · Mjukvaru- och webbutveckling · Mjukvaruutveckling · Molnteknik · Datavetenskap
I korthet
The AI Infrastructure Engineer (L4) is an enterprise-level architectural leadership role responsible for designing, governing, and scaling next-generation GPU/accelerator platforms for AI/ML workloads. This domain leader drives roadmap, standards, and cross-functional alignment, optimizing Linux systems and large-scale distributed training/inference architectures.
Ansvarsområden
- Define and own end-to-end architecture for enterprise AI infrastructure platforms (GPU, storage, network, orchestration).
- Establish reference architectures, design standards, and best practices for AI/ML platforms across hybrid/cloud environments.
- Lead platform modernization (AI-native infrastructure, automation-first, GPUaaS models).
- Drive platform scalability strategy for multi-cluster, multi-region deployments.
- Optimize Linux systems (Ubuntu, RHEL, Rocky) for AI/HPC workloads through NUMA, kernel, and clock tuning.
- Define long-term roadmap for AI infrastructure aligned with business and AI/GenAI adoption strategy.
- Evaluate and onboard emerging GPU/accelerator technologies (NVIDIA, AMD, TPU, specialized AI hardware).
- Lead vendor strategy and partnerships (OEMs, cloud providers, NVIDIA ecosystem, etc.).
- Provide strategic advisory to leadership on AI infrastructure investments and scaling decisions.
- Lead optimization of large-scale distributed training and inference architectures.
- Establish best practices for LLM training/inference platforms (vLLM, Triton, TensorRT-LLM, DeepSpeed).
Krav
- Bachelor’s/master’s degree in computer science, Engineering, or related field.
- 12–18 years of overall infrastructure/platform engineering experience.
- 6–10 years in AI/ML infrastructure and distributed systems at scale.
- Proven experience in architecting enterprise AI platforms (on-prem + cloud + hybrid).
- Deep expertise in Kubernetes at scale.
- Deep expertise in GPU infrastructure (NVIDIA ecosystem).
- Deep expertise in HPC and distributed training frameworks.
- Strong exposure to GenAI / LLM workloads and phantomization.
- Demonstrated experience in leading large programs.
- Demonstrated experience in defining architecture & strategy.
- Demonstrated experience in customer-facing solutioning.
- Knowledge on Pacemaker and devops.
- Linux, Pacemaker.
- Very good understanding on Hardware, Linux operating system, performance management, network.
- Good Knowledge on Kubernetes, Virtualization, Ansible, Red hat Satellite.
- Knowledge on Hyperscalers like AWS/Azure/GCP.
Önskade kvalifikationer
- NVIDIA Certified AI Infrastructure (Associate/Professional).
- CKA / CKS (Kubernetes).
- AWS / Azure Architect (Professional level preferred).
- Relevant Red Hat Certifications (E.G., Red Hat Certified Engineer Rhce, Red Hat Certified Specialist In Ansible Automation) Are Optional But Valuable.
Förmåner
- Supercharge your potential.
- Find your career.
- Find your spark.
- A place that knows that helping its customers stay on top starts by putting its people first.
#AI Infrastructure#ML Infrastructure#GPU Platforms#HPC#Kubernetes#Linux#Red Hat#Satellite#Cluster#DevOps#Nvidia#Nvidia GPU#Enterprise Architecture#Scalability#Performance Optimization#GenAI#LLM