Software Engineering Manager, AI/ML Infrastructure and Performance Engineering

Technology, Data & Digital · Data, AI & Analytics · Machine Learning · Software Engineering · Data Engineering

In short

As a Software Engineering Manager, you will lead teams focused on AI/ML Infrastructure and Performance Engineering, driving continuous improvements in ML software/hardware stacks. You will provide technical leadership, manage engineers, and contribute to product strategy to optimize ML systems and applications at scale.

Responsibilities

  • Learn and build an intuitive understanding of existing data collection, analysis, and visualization workflows with deep introspection across Frameworks, Accelerated Linear Algebra (XLA) and runtime stack.
  • Support new and exciting ML paradigms (such as horizontal scaling for upcoming TPU chips) by making contributions across the end to end stack and analysis tools.
  • Partner with ML Stack leads to understand model optimization use cases and bring debugging to feel native in 3P environments (VSCode, Cursor, Grafana, etc).
  • Work with OSS ML inference frameworks such as vLLM, SGLang to provide insights in Xprof about performance improvement opportunities.
  • Partner with other teams that own various parts of the ML stack to understand performance optimization use cases.
  • Work with OSS ML inference frameworks such as TorchTPU, vLLM, SGLang to provide insights into performance bottlenecks.

Requirements

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience with software engineering, machine learning infrastructure, computer architecture, distributed computing, people management, debugging tools, communication.
  • Experience with people management, and building infrastructure to improve performance of Machine Learning (ML) systems and applications.

Desired Qualifications

  • Experience with Performance Optimization, Graphics Processing Unit (GPU) Programming, High Performance Computing, Large Language Model, Open Source Contributor.
  • Experience with the Machine Learning infra and frameworks and hands-on experience with GPU or Tensor Processing Unit (TPU) performance analysis.
  • Experience in building agentic workflows for performance debugging and optimization. Ability to generate ideas and resolve ambiguity.
  • Hands-on experience with Machine Learning (ML) frameworks such as TensorFlow, JAX, and PyTorch, Keras.
  • Experience with ML Inference frameworks such as vLLM, SG Lang, Pathways and experience in open-source software development, including experience in releasing and supporting open-source projects.

Skills

TensorFlowPyTorchKerasJAXvLLMSGLang
#AI/ML#Infrastructure#Performance Engineering#Software Engineering Management
Google Logo

About Google

Innovative Full-Stack Developer | Building Scalable Solutions with Node.js & React

At Google, our mission is to organize the world's information and make it universally accessible and useful. From Search to AI, we develop products and services that improve the lives of billions, empowering individuals and businesses globally. We're driven by a passion for innovation, a commitment to solving complex problems, and a belief in the transformative power of technology. Since our founding in 1998, we've remained dedicated to creating a more helpful and connected world.

Google Logo

Company

Google

Job Posted

1 day ago

Employment Type

Full Time

Work mode

On Site

Experience Level

Senior

Locations

Bengaluru, India

Qualification

Bachelor

Applicants

Be an early applicant