Engineer-ARM Neon,ARM SVE,Performance optimization

Technology, Data & Digital · Software & Web Development · Software Engineering · Embedded Systems

In short

Qualcomm is seeking an Engineer specializing in ARM Neon, ARM SVE, and performance optimization for CPU software & hardware co-design in Bangalore, India. This role focuses on optimizing machine learning workloads on next-generation QMX architectures through workload characterization, simulation, kernel optimization, and providing architectural feedback.

Responsibilities

  • Identify and prioritize critical ML use cases and models for CPU-centric execution.
  • Analyze workload characteristics including compute intensity, memory bandwidth, cache behavior, and parallelism.
  • Generate detailed execution traces for ML workloads using QEMU or equivalent simulators.
  • Capture instruction-level execution behavior and extract performance counters and bottlenecks.
  • Identify system bottlenecks across CPU pipelines, memory hierarchy, and instruction utilization.
  • Optimize critical hotspots through kernel-level tuning, algorithmic improvements, and data layout optimizations.
  • Collaborate with CPU architecture and design teams to provide data-driven insights and propose architectural enhancements.
  • Influence next-generation CPU features in compute units, vector/SIMD extensions (e.g., QMX), and memory subsystems.
  • Design and implement highly optimized ML kernels and libraries for QMX architecture.
  • Develop kernels for GEMM, convolution, attention, activation functions, etc.
  • Optimize CPU-centric ML benchmarks such as Geekbench AI and internal benchmarking suites.
  • Establish performance baselines and track improvements across hardware generations.

Requirements

  • Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
  • OR Master's degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience.
  • OR PhD in Engineering, Information Systems, Computer Science, or related field.
  • 2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.
  • Strong background in Computer Architecture / Systems Programming.
  • Strong background in Machine Learning fundamentals.
  • Proficiency in C/C++ (mandatory).
  • Experience with performance profiling, benchmarking, and optimization.

Desired Qualifications

  • Experience with QEMU or equivalent simulators.
  • Experience with ML kernel development (GEMM, convolution, attention).
  • Knowledge of CPU architecture (pipelines, caching, SIMD/vector extensions such as NEON, SVE, QMX).
  • Familiarity with ML frameworks and inference stacks.
  • Experience with low-level optimization: Intrinsics, assembly, memory and cache tuning.

Benefits

  • Work on next-generation CPU architectures (QMX).
  • Directly influence hardware design through real workload insights.
  • Solve end-to-end ML performance challenges (model → kernel → silicon).
  • Collaborate with top architecture, systems, and AI teams.
  • High-impact role with visibility across product and research roadmaps.
#ARM Neon#ARM SVE#Performance optimization#CPU architecture#Machine Learning#System-level performance#QMX#ML kernels#SIMD
Qualcomm Logo

Company

Qualcomm

Job Posted

1 month ago

Employment Type

Full Time

WorkMode

On Site

Experience Level

Mid-Senior

Locations

Bangalore, India

Qualification

Bachelor

Applicants

Be an early applicant