Sr. AI Accelerator Tools Development Engineer
Technology, Data & Digital · Software & Web Development · Embedded Systems · Software Engineering
In short
Microsoft Silicon, Cloud Hardware, and Infrastructure Engineering is looking for a Sr. AI Accelerator Tools Development Engineer to lead the development of next-generation stress, validation, and performance tooling for MAIA AI accelerator platforms. This role involves building software frameworks and workloads that exercise the entire AI stack, contributing to platform bring-up, qualification, and reliability.
Responsibilities
- Design and develop scalable stress, performance, and validation frameworks for MAIA AI accelerator platforms.
- Build workload generation infrastructure capable of exercising compute, memory, interconnect, networking, storage, and system-level resources.
- Develop reusable stress tools using PyTorch, Triton, Python, C++, and custom MAIA SDKs.
- Create synthetic and production-inspired workloads that model training and inference behaviors observed in large-scale AI deployments.
- Build automated infrastructure for workload deployment, orchestration, telemetry collection, and result analysis.
- Develop and optimize kernels targeting custom AI accelerators.
- Create GEMM, attention, collective communication, and memory intensive stress workloads.
- Analyze execution behavior across the hardware-software stack and identify bottlenecks impacting utilization and performance.
- Collaborate with compiler and runtime teams to improve workload efficiency and hardware utilization.
- Develop tooling that integrates with MAIA compiler pipelines, SDKs, runtime environments, and performance analysis tools.
- Understand and debug compiler output, generated kernels, scheduling decisions, and execution behavior.
- Build automation around model compilation, kernel validation, regression testing, and workload portability.
- Partner with compiler teams to validate new compiler features and workload optimization strategies.
- Design workload suites for platform bring-up, qualification, and reliability testing.
- Build comprehensive regression infrastructure supporting silicon, firmware, system software, and platform releases.
- Develop automated validation tools capable of identifying correctness, performance, thermal, power, and stability issues.
- Enable platform readiness through scalable validation methodologies and continuous regression testing.
- Characterize system performance across compute, networking, memory, and storage subsystems.
- Develop benchmarking methodologies and performance dashboards.
- Adapt and optimize industry-standard workloads including HPL/HPC benchmarks, LLM training workloads, Transformer-based inference workloads, Collective communication benchmarks, and AI framework benchmark suites.
- Drive root-cause analysis and optimization initiatives across the stack.
- Improve developer productivity through automation, CI/CD integration, diagnostics, and debugging infrastructure.
- Build reusable tooling for workload generation, failure triage, telemetry analysis, and reporting.
- Develop dashboards and automated workflows for large-scale validation environments.
- Partner with engineering teams to convert recurring validation challenges into durable tooling solutions.
Requirements
- Doctorate in Electrical Engineering, Computer Engineering, Computer Science, or related field AND 1+ year(s) technical engineering experience OR Master's Degree in Electrical Engineering, Computer Engineering, Computer Science, or related field AND 4+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Computer Science, or related field AND 5+ years technical engineering experience OR equivalent experience.
- 4+ years of experience with programming skills in C/C++ and Python, with experience designing and developing production-quality software systems, frameworks, runtimes, or infrastructure.
- 4+ years of experience developing, debugging, or optimizing AI workloads for GPUs, AI accelerators, or HPC systems, including experience with PyTorch or comparable AI frameworks and compute-intensive kernels.
- 4+ years of experience with accelerator architecture and performance optimization, including compute engines, memory hierarchy, interconnects, runtime systems, performance profiling, bottleneck analysis, and debugging across hardware/software boundaries.
- Ability to meet Microsoft, customer and/or government security screening requirements.
Desired Qualifications
- Experience with custom AI accelerator SDKs, compiler ecosystems, kernel generation frameworks, or code-generation pipelines, including technologies such as LLVM, MLIR, or Triton.
- Experience developing or optimizing AI training and inference workloads, including LLMs and other large-scale AI models.
- Experience with collective communication libraries, high-performance networking, distributed AI systems, or performance characterization of large-scale AI clusters.
- Experience with CI/CD systems, containerized environments, automated testing, or cloud-scale validation infrastructure.
- Experience working with silicon bring-up or post-silicon validation teams.
Benefits
- Certain roles may be eligible for benefits and other compensation.
- Find additional benefits and pay information here: https://careers.microsoft.com/us/en/us-corporate-pay
#azure#MAIA#AI/ML#firmware engineering#hardware engineering#silicon