Senior Site Reliability Engineer - Azure Storage

Technology, Data & Digital · IT Infrastructure & Security · Site Reliability Engineering · Software Engineering · QA & Testing

In short

Microsoft Azure Storage is seeking a Senior Site Reliability Engineer to lead the qualification and performance validation of next-generation storage infrastructure. This remote role involves driving strategy, developing automation, and resolving complex issues across hardware, software, and firmware to ensure the reliability and scalability of Azure Storage.

Responsibilities

  • Lead end-to-end qualification and production readiness validation for new Azure Storage hardware SKUs, platforms, SSDs, networking devices, and firmware releases.
  • Drive qualification strategy, test planning, execution, and sign-off recommendations for Azure Storage deployments.
  • Serve as a technical escalation point for qualification failures, performance regressions, reliability issues, and production-readiness concerns impacting Azure Storage deployments.
  • Own complex investigations across software, hardware, firmware, and platform stacks to identify and resolve qualification blockers.
  • Design and develop scalable automation frameworks that improve qualification coverage, efficiency, quality, and engineering productivity.
  • Analyze large-scale telemetry, benchmark data, and fleet health signals to identify performance trends, bottlenecks, risks, and optimization opportunities.
  • Monitor qualification pipelines and proactively identify risks impacting production onboarding, deployment schedules, and service readiness.
  • Perform root cause analysis for hardware, firmware, platform, and software issues discovered during qualification and drive corrective actions to closure.
  • Validate SKU readiness against Azure Storage performance, reliability, scalability, and operational requirements.
  • Partner with Storage, Compute, Networking, Platform, and vendor engineering teams to resolve critical qualification, reliability, and performance issues.
  • Influence qualification standards, test methodologies, performance thresholds, and operational best practices across Azure Storage NPI programs.

Requirements

  • Bachelor's Degree in Computer Science, Engineering, or related field AND 4+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience.
  • Ability to meet Microsoft, customer and/or government security screening requirements.

Desired Qualifications

  • Bachelor Degree in Computer Science, Engineering, or related field AND 6+ years technical experience in service engineering, performance engineering, or site reliability engineering.
  • 6+ years technical experience working with large-scale cloud or distributed systems.

Benefits

  • Base pay range of USD $119,800 - $234,700 per year (US), with a different range applicable to San Francisco Bay area and New York City metropolitan area (USD $160,200 - $261,000 per year).
  • Certain roles may be eligible for benefits and other compensation.
#cloud#storage#reliability#performance#automation#hardware#firmware#networking#site reliability engineering#azure
Microsoft Logo

Company

Microsoft

Job Posted

2 weeks ago

Employment Type

Full Time

WorkMode

Remote

Experience Level

Senior

Locations

United States

Qualification

Bachelor

Applicants

Be an early applicant