AI Model Optimization Architect
Technology, Data & Digital · Software & Web Development · Software Engineering · Machine Learning · Data Science
In short
Qualcomm is seeking an AI Model Optimization Architect in Cork, Ireland, to lead model transformation and optimization for large-scale AI models on their inference accelerators. This role involves deep collaboration with compiler and performance teams to enhance throughput, latency, and quality for LLMs, VLMs, and diffusion models.
Responsibilities
- Architect and deliver model optimization strategies transforming PyTorch models for efficient inference on Qualcomm accelerators.
- Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations.
- Design and implement fusion kernels using DSL based approaches (e.g., Triton), enabling fused operations and performance critical algorithmic rewrites.
- Partner deeply with compiler, performance, and accuracy teams to co-design lowering strategies, kernel fusion, layout decisions, and runtime integration.
- Profile and optimize LLM/VLM/diffusion inference for throughput and latency across batch sizes, sequence lengths, and serving modes.
- Own transformer specific optimizations including KVcache management, decoding behavior, and long context performance.
- Enable and optimize continuous batching (dynamic/iteration-level scheduling), understanding its impact on memory, scheduling, and tail latency.
- Architect and scale distributed inference strategies (e.g., sharding and parallelism) across multi-core and multi-device systems.
- Establish reusable approaches to scale model optimizations to new hardware architectures, creating robust patterns and tooling.
- Debug complex performance or stability issues to root cause and drive production ready solutions.
Requirements
- Expert level expertise in PyTorch and inference focused model optimization; strong Python engineering skills.
- Hands on experience with torch.compile / TorchDynamo or related graph capture and compilation workflows.
- Deep understanding of transformer architectures, attention mechanisms, MoEs, and performance trade-offs.
- Practical experience with KVcache behavior, serving time optimizations, and memory/performance tradeoffs.
- Strong foundation in computer architecture, ML accelerators, and distributed systems.
- Proven ability to lead cross-functional technical efforts and influence design decisions.
- MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering, or equivalent experience.
- Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 6+ years of Software Engineering or related work experience.
- OR Master's degree in Engineering, Information Systems, Computer Science, or related field and 5+ years of Software Engineering or related work experience.
- OR PhD in Engineering, Information Systems, Computer Science, or related field and 4+ years of Software Engineering or related work experience.
- 3+ years of work experience with Programming Language such as C, C++, Java, Python, etc.
Desired Qualifications
- Experience developing fusion kernels using Triton or similar DSLs, and collaborating with ML compiler teams.
- Familiarity with LLM serving stacks and continuous batching systems.
- Background in numerical methods, performance/accuracy trade-off analysis, or evaluation frameworks.
Benefits
- Salary, stock and performance related bonus
- Maternity/Paternity Leave
- Employee stock purchase scheme
- Matching pension scheme
- Education Assistance
- Relocation and immigration support (if needed)
- Life, Medical, Income and Travel Insurance
- Subsidised memberships for physical and mental well-being
- Bicycle purchase scheme
- Employee run clubs, including, running, football, chess, badminton + many more
#AI#Machine Learning#Optimization#Architecture#LLM#VLM#Diffusion Models#Multimodal Models#Inference Accelerators#PyTorch#ONNX#Triton#Compiler#Performance#Quality#Throughput#Latency#Memory#Transformer Architectures#KV Cache#Continuous Batching#Distributed Inference#Sharding#Parallelism#Cloud AI