Sr. AI Accelerator Tools Development Engineer

Technology, Data & Digital · Software & Web Development · Embedded Systems · Software Engineering

In short

Microsoft Silicon, Cloud Hardware, and Infrastructure Engineering is looking for a Sr. AI Accelerator Tools Development Engineer to lead the development of next-generation stress, validation, and performance tooling for MAIA AI accelerator platforms. This role involves building software frameworks and workloads that exercise the entire AI stack, contributing to platform bring-up, qualification, and reliability.

Responsibilities

  • Design and develop scalable stress, performance, and validation frameworks for MAIA AI accelerator platforms.
  • Build workload generation infrastructure capable of exercising compute, memory, interconnect, networking, storage, and system-level resources.
  • Develop reusable stress tools using PyTorch, Triton, Python, C++, and custom MAIA SDKs.
  • Create synthetic and production-inspired workloads that model training and inference behaviors observed in large-scale AI deployments.
  • Build automated infrastructure for workload deployment, orchestration, telemetry collection, and result analysis.
  • Develop and optimize kernels targeting custom AI accelerators.
  • Create GEMM, attention, collective communication, and memory intensive stress workloads.
  • Analyze execution behavior across the hardware-software stack and identify bottlenecks impacting utilization and performance.
  • Collaborate with compiler and runtime teams to improve workload efficiency and hardware utilization.
  • Develop tooling that integrates with MAIA compiler pipelines, SDKs, runtime environments, and performance analysis tools.
  • Understand and debug compiler output, generated kernels, scheduling decisions, and execution behavior.
  • Build automation around model compilation, kernel validation, regression testing, and workload portability.
  • Partner with compiler teams to validate new compiler features and workload optimization strategies.
  • Design workload suites for platform bring-up, qualification, and reliability testing.
  • Build comprehensive regression infrastructure supporting silicon, firmware, system software, and platform releases.
  • Develop automated validation tools capable of identifying correctness, performance, thermal, power, and stability issues.
  • Enable platform readiness through scalable validation methodologies and continuous regression testing.
  • Characterize system performance across compute, networking, memory, and storage subsystems.
  • Develop benchmarking methodologies and performance dashboards.
  • Adapt and optimize industry-standard workloads including HPL/HPC benchmarks, LLM training workloads, Transformer-based inference workloads, Collective communication benchmarks, and AI framework benchmark suites.
  • Drive root-cause analysis and optimization initiatives across the stack.
  • Improve developer productivity through automation, CI/CD integration, diagnostics, and debugging infrastructure.
  • Build reusable tooling for workload generation, failure triage, telemetry analysis, and reporting.
  • Develop dashboards and automated workflows for large-scale validation environments.
  • Partner with engineering teams to convert recurring validation challenges into durable tooling solutions.

Requirements

  • Doctorate in Electrical Engineering, Computer Engineering, Computer Science, or related field AND 1+ year(s) technical engineering experience OR Master's Degree in Electrical Engineering, Computer Engineering, Computer Science, or related field AND 4+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Computer Science, or related field AND 5+ years technical engineering experience OR equivalent experience.
  • 4+ years of experience with programming skills in C/C++ and Python, with experience designing and developing production-quality software systems, frameworks, runtimes, or infrastructure.
  • 4+ years of experience developing, debugging, or optimizing AI workloads for GPUs, AI accelerators, or HPC systems, including experience with PyTorch or comparable AI frameworks and compute-intensive kernels.
  • 4+ years of experience with accelerator architecture and performance optimization, including compute engines, memory hierarchy, interconnects, runtime systems, performance profiling, bottleneck analysis, and debugging across hardware/software boundaries.
  • Ability to meet Microsoft, customer and/or government security screening requirements.

Desired Qualifications

  • Experience with custom AI accelerator SDKs, compiler ecosystems, kernel generation frameworks, or code-generation pipelines, including technologies such as LLVM, MLIR, or Triton.
  • Experience developing or optimizing AI training and inference workloads, including LLMs and other large-scale AI models.
  • Experience with collective communication libraries, high-performance networking, distributed AI systems, or performance characterization of large-scale AI clusters.
  • Experience with CI/CD systems, containerized environments, automated testing, or cloud-scale validation infrastructure.
  • Experience working with silicon bring-up or post-silicon validation teams.

Benefits

  • Certain roles may be eligible for benefits and other compensation.
  • Find additional benefits and pay information here: https://careers.microsoft.com/us/en/us-corporate-pay
#azure#MAIA#AI/ML#firmware engineering#hardware engineering#silicon
Microsoft Logo

Company

Microsoft

Job Posted

1 day ago

Employment Type

Full Time

WorkMode

Hybrid

Experience Level

Senior

Locations

United States

Qualification

Doctoral, Master, Bachelor

Applicants

Be an early applicant