LLM Serving Engineer (Cloud AI Engineering), Senior / Staff Engineer - Cork, Ireland
Technology, Data & Digital · Software & Web Development · Software Engineering · Machine Learning
In short
Qualcomm is seeking experienced LLM Serving Engineers in Cork, Ireland, to develop and deploy scalable LLM inference platforms. This role involves working with advanced inference techniques, contributing to LLM Serving packages, and collaborating with customers and internal teams at the forefront of GenAI research and development.
Responsibilities
- Build scalable LLM inference platforms using techniques like disaggregated serving, KV-Cache management, advanced parallelism, speculative algorithms, model optimization, and specialized kernels.
- Contribute to the development of LLM Serving packages (e.g., vLLM, SGLang, TGI, Triton-Inference server, Dynamo, LLM-d).
- Collaborate with customers and internal compiler, firmware, and platform teams to drive solutions.
- Understand advanced algorithms (e.g., attention mechanisms, MoEs) and numerics for optimization opportunities in GenAI.
- Drive efficient serving through smart autoscaling, load balancing, and routing.
- Engage with open-source serving communities to evolve the framework.
Requirements
- Hands-on experience in one or more LLM serving/Orchestration packages (Triton-Inference Server, vLLM, SGLang, Ollama, llm-d, KServe, LMCache, MoonCake).
- Deep understanding of foundational LLMs, VLMs, SLMs, and transformer-based architectures.
- Strong experience in developing language models using PyTorch.
- Strong computer science fundamentals: algorithms, data structures, parallel and distributed programming.
- Understanding of computer architecture, ML accelerators, in-memory processing, and distributed systems.
- Strong Python development skills for large-scale projects.
- Experience in analyzing, profiling, and optimizing deep learning workloads.
- Proactive learning about the latest inference optimization techniques.
- Excellent communication and problem-solving skills.
- MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering.
- Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 4+ years of Software Engineering or related work experience.
- Master's degree in Engineering, Information Systems, Computer Science, or related field and 3+ years of Software Engineering or related work experience.
- PhD in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
- 2+ years of work experience with Programming Languages such as C, C++, Java, Python, etc.
Desired Qualifications
- Open-source contribution to any GenAI package.
- Experience architecting and developing large-scale distributed systems.
- High-level kernel design experience (PyTorch, CUDA, Triton).
- Knowledge of torch.compile or torchDynamo.
Benefits
- Salary, stock, and performance-related bonus
- Maternity/Paternity Leave
- Employee stock purchase scheme
- Matching pension scheme
- Education Assistance
- Relocation and immigration support (if needed)
- Life, Medical, Income, and Travel Insurance
- Subsidised memberships for physical and mental well-being
- Bicycle purchase scheme
- Employee run clubs (running, football, chess, badminton, etc.)
#Cloud AI#LLM Serving#Inference Acceleration#GenAI#Machine Learning#Software Engineering