Associate Engineer
In short
Qualcomm is hiring an Associate Engineer focused on AI/ML performance. This role involves analyzing and optimizing AI workloads across NPUs and GPUs, focusing on inference performance, latency, and power efficiency. The position requires a Bachelor's degree and experience with embedded software, system performance, and machine learning fundamentals.
Responsibilities
- Analyze and optimize AI/ML workload performance on NSP, NPU, and GPU across single and multi‑accelerator systems.
- Identify bottlenecks across compute, memory bandwidth/latency, NoC, DMA, SMMU, and scheduling.
- Drive optimizations for latency‑critical inference, high‑throughput pipelines, and concurrent multi‑VM execution.
- Perform roofline, utilization, and memory access pattern analysis using hardware counters.
- Benchmark and optimize CNNs, Transformers, VLMs, and diffusion models for real‑world and customer use cases.
- Optimize model execution using operator fusion, graph partitioning, quantization, mixed precision, tiling, and batching.
- Collaborate with system software teams on drivers, runtimes, memory allocation, power, thermal, and QoS constraints.
- Analyze performance under concurrent, virtualized, and safety‑critical or real‑time environments.
- Build and use benchmarks, micro‑benchmarks, profilers, and regression tools to derive actionable insights.
- Partner with architecture, compiler, SDK, and AI framework teams to influence HW/SW design and resolve performance escalations.
Requirements
- Bachelor's degree in Computer Science Engineering, Information Systems, or related field.
- 1+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
- Good software development skills with exposure to AI coding tools.
- Excellent analytical, development, and problem-solving skills.
- Experience in embedded software (Linux, Android, RTOS).
- Experience in system performance domain with focus on neural accelerators and GPU.
- Knowledge on CPU, NPU, GPU, NOC, DDR architecture.
- Strong understanding of Machine Learning fundamentals.
- Knowledge in neural network quantization, compression, pruning algorithms, deep learning kernel/compiler optimization.
- Strong communication skills.
#AI/ML#Performance Engineering#SoC#NPU#GPU#Deep Learning#Inference#Latency#Throughput#Power Efficiency#Embedded Software#System Performance#Neural Accelerators#Architecture#Compiler#Runtime#Frameworks