PE- CPU Software & Hardware Co-Design Engineer (ML Systems)
Technology, Data & Digital · Software & Web Development · Software Engineering · Embedded Systems
In short
Qualcomm is looking for a CPU Software & Hardware Co-Design Engineer specializing in ML Systems in Bangalore, India. This role focuses on CPU software-hardware co-design for next-generation QMX architectures, involving workload characterization, simulation, kernel optimization, and providing architectural insights for future CPU designs. The ideal candidate will work across the full stack to enable efficient execution of ML workloads on CPU platforms.
Responsibilities
- Identify and prioritize critical ML use cases and models for CPU-centric execution.
- Analyze workload characteristics including compute intensity, memory bandwidth, cache behavior, parallelism and dataflow patterns.
- Generate detailed execution traces for ML workloads using QEMU or equivalent simulators.
- Develop tooling to capture instruction-level execution behavior and extract performance counters.
- Identify system bottlenecks and optimize critical hotspots through kernel-level tuning, algorithmic improvements, and data layout optimizations.
- Collaborate with CPU architecture and design teams to provide data-driven insights and propose architectural enhancements.
- Design and implement highly optimized ML kernels and libraries for QMX architecture.
- Enable integration with open-source ML frameworks like PyTorch, ONNX, XNNPACK, MLAS.
- Optimize CPU-centric ML benchmarks and establish performance baselines.
- Perform competitive analysis and performance positioning.
Requirements
- Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 8+ years of Software Engineering experience.
- Master's degree in Engineering, Information Systems, Computer Science, or related field and 7+ years of Software Engineering experience.
- PhD in Engineering, Information Systems, Computer Science, or related field and 6+ years of Software Engineering experience.
- 4+ years of work experience with Programming Language such as C, C++, Java, Python, etc.
- Strong background in Computer Architecture / Systems Programming.
- Strong background in Machine Learning fundamentals.
- Proficiency in C/C++ (mandatory).
- Experience with performance profiling, benchmarking, and optimization.
Desired Qualifications
- Experience with QEMU or equivalent simulators.
- Experience with ML kernel development (GEMM, convolution, attention).
- Knowledge of CPU architecture (pipelines, caching, SIMD/vector extensions such as NEON, SVE, QMX).
- Familiarity with ML frameworks and inference stacks.
- Experience with low-level optimization: intrinsics, assembly, memory and cache tuning.
Benefits
- Work on next-generation CPU architectures (QMX).
- Directly influence hardware design through real workload insights.
- Solve end-to-end ML performance challenges (model → kernel → silicon).
- Collaborate with top architecture, systems, and AI teams.
- High-impact role with visibility across product and research roadmaps.
#CPU architecture#machine learning#system-level performance optimization#co-design#QMX#workload characterization#simulation#kernel optimization#architectural feedback#ML models#CPU platforms#LLMs#vision#speech#recommender systems#compute intensity#memory bandwidth#cache behavior#parallelism#dataflow#QEMU#instruction-level execution#performance counters#bottlenecks#CPU pipelines#instruction utilization#kernel-level tuning#algorithmic improvements#data layout#memory optimizations#compute units#Vector/SIMD extensions#NEON#SVE#memory subsystems#GEMM#convolution#attention#activation functions#inference stacks#intrinsics#assembly#Geekbench AI