Software Engineer 2

Technology, Data & Digital · Software & Web Development · Software Engineering · Cloud Engineering · DevOps

In short

Microsoft is looking for a Software Engineer 2 to design and build scalable AKS/Kubernetes infrastructure for GPU and AI accelerator platforms. The role involves working on cluster lifecycle, node management, deployment, and ensuring reliability at cloud scale. This position is based in India and requires a Bachelor's degree in Computer Science or related field, with 4+ years of experience.

Responsibilities

  • Design and develop scalable AKS/Kubernetes infrastructure for GPU and AI accelerator environments.
  • Build services and automation for cluster provisioning, configuration, upgrades and lifecycle management.
  • Develop solutions for node lifecycle, health monitoring, failure detection and automated recovery.
  • Improve infrastructure scalability, reliability and availability across large accelerator fleets.
  • Build Kubernetes integrations for accelerator discovery, resource management, scheduling and workload enablement.
  • Develop reliable software and automation for deployment, configuration and management of accelerator infrastructure.
  • Improve telemetry, observability, diagnostics and operational readiness of distributed infrastructure.
  • Diagnose complex issues across Kubernetes, containers, Linux, networking and accelerator infrastructure.
  • Collaborate with Control Plane, systems software, hardware and platform teams to deliver end-to-end solutions.
  • Contribute to architecture/design reviews, engineering practices, code quality and mentoring of other engineers.
  • Apply AI-assisted engineering practices across design, coding, testing, debugging, code reviews and documentation to improve engineering velocity and quality.
  • Use AI-assisted workflows to accelerate code comprehension, troubleshooting, root-cause analysis, test development and infrastructure automation.
  • Leverage AI to accelerate learning and build deeper AKS/Kubernetes, distributed systems and AI infrastructure domain competency.
  • Identify repetitive engineering and operational workflows that can be simplified or automated using AI-enabled engineering approaches.
  • Share reusable AI-assisted engineering practices and technical knowledge to improve team productivity and engineering competency.

Requirements

  • Bachelor’s Degree in Computer Science, Computer Engineering or related technical discipline, or equivalent experience with 4+ years of industry relevant experience.
  • Software development skills in Go, C++, C#, Python or similar languages.
  • Experience building distributed systems, cloud infrastructure or platform services.
  • Hands-on experience with Kubernetes, containerisation and cloud-native technologies.
  • Understanding of distributed-systems concepts including scalability, concurrency, state management, resiliency and failure recovery.
  • Experience developing reliable production software and debugging complex distributed systems.
  • Design, problem-solving and cross-team collaboration skills.
  • Demonstrated ability or experience using AI-assisted software engineering tools and workflows to improve development effectiveness, debugging, automation and software quality.
  • Ability to rapidly develop expertise in complex infrastructure technologies and apply that knowledge to production engineering problems.

Desired Qualifications

  • Experience with Azure Kubernetes Service (AKS) or large-scale Kubernetes environments.
  • Experience with cluster/node provisioning, Kubernetes controllers/operators, scheduling or resource management.
  • Experience operating and scaling Kubernetes-based production infrastructure.
  • Experience with GPU, AI accelerator or heterogeneous compute infrastructure.
  • Experience with Linux, containers, networking and host-level system software.
  • Knowledge of accelerator virtualisation and resource isolation concepts.
  • Experience with infrastructure automation, CI/CD and configuration/deployment systems.
  • Experience with telemetry, metrics, distributed tracing, observability and production diagnostics.
  • Experience applying AI-assisted development to code generation, code comprehension, testing, troubleshooting, documentation and engineering automation.
  • Experience using AI-enabled workflows to investigate complex distributed-system and infrastructure failures.
  • Experience leveraging AI to automate repetitive infrastructure engineering and operational tasks.
  • Ability to use AI-assisted learning alongside engineering fundamentals to accelerate development of Kubernetes, cloud infrastructure and accelerator-domain expertise.
  • Experience sharing reusable AI-assisted engineering workflows and practices across an engineering team.

Skills

KubernetesGoC++C#PythonLinuxDockerAKS
#AI#Kubernetes#Cloud#Infrastructure#GPU#Accelerator
Microsoft Logo

About Microsoft

Empowering every person and organization on the planet to achieve more

Every company has a mission. What's ours? To empower every person and every organization to achieve more. We believe technology can and should be a force for good and that meaningful innovation contributes to a brighter world in the future and today. Our culture doesn’t just encourage curiosity; it embraces it. Each day we make progress together by showing up as our authentic selves. We show up with a learn-it-all mentality. We show up cheering on others, knowing their success doesn't diminish our own. We show up every day open to learning our own biases, changing our behavior, and inviting…

Microsoft Logo

Company

Microsoft

Job Posted

1 day ago

Employment Type

Full Time

Work mode

Hybrid

Experience Level

Associate

Locations

India

Qualification

Bachelor

Applicants

Be an early applicant