Sr. Software Engineer
Technology, Data & Digital · Software & Web Development · Software Engineering
In short
Microsoft Azure AI Inference platform is looking for a Sr. Software Engineer to design, optimize, and scale inference systems for cutting-edge AI models, including LLMs and GenAI from providers like OpenAI. This role involves working on high-throughput, low-latency environments and influencing the overall product and platform capabilities.
Responsibilities
- Design and implement core inference infrastructure for serving frontier AI models in production.
- Identify and drive improvements to end-to-end inference performance and efficiency of state-of-the-art LLMs and GenAI models from OpenAI, Anthropic and xAI hosted on AI Foundary.
- Design and implement efficient load scheduling and balancing strategies, by leveraging key insights and features of the model and workload.
- Scale the platform to support the growing inferencing demand and maintain high availability.
- Deliver critical capabilities required to serve the latest and greatest Gen AI models such as GPT5, Realtime audio, Sora, and enable fast time to market for them.
- Drive generic features to cater to the needs of customers such as GitHub, M365, Microsoft AI and third-party companies.
- Collaborate with partners both internal and external.
- Embody Microsoft's Culture and Values.
Requirements
- Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, or Java OR equivalent experience.
- Ability to meet Microsoft, customer and/or government security screening requirements.
- Pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.
Desired Qualifications
- Technical background and solid foundation in software engineering principles, distributed computing and architecture.
- Experience working on high scale, reliable online systems and experience in real-time online services with low latency and high throughput.
- Experience working with L7 network proxies and gateways.
- Knowledge in Network architecture and concepts (HTTP and TCP Protocols, Authentication and Sessions etc).
- Knowledge and experience in OSS, Docker, Kubernetes, C++, Golang, or equivalent programming languages.
- Cross-team collaboration skills and the desire to collaborate in a team of researchers and developers.
Benefits
- The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year.
- In the San Francisco Bay area and New York City metropolitan area, the base pay range for this role is USD $160,200 - $261,000 per year.
- Certain roles may be eligible for benefits and other compensation.
#AI#Azure#Inference#LLM#GenAI#OpenAI#Cloud