Cloud Network Engineering IC4
Teknik, data och digitalt · IT-infrastruktur och säkerhet · Molnteknik · Mjukvaruutveckling
I korthet
The HPC/AI team is building the next-generation distributed AI supercomputer, and you will help deliver and operate the network infrastructure. As a Senior Cloud Network Engineer, you will manage AI network fabrics critical for ultra-low latency and high throughput in distributed AI workloads, working at the intersection of AI supercomputing and large-scale networking.
Ansvarsområden
- Develop technical roadmap and servicing protocols/procedures for BackEnd Network infrastructure.
- Create and propose technical solutions for complex systems design or process issues.
- Provide technical expertise for root cause analysis of service-impacting incidents and implement prevention strategies.
- Apply influence and negotiation skills to impact decisions based on TCO and risk.
- Drive continuous improvements and automation for networking services to manage scale and optimize efforts.
- Serve as a Subject Matter Expert in areas like network operating systems, network automation in Python, BGP, cloud-scale operations, MPLS, large-scale testing, OpenConfig, or Streaming Telemetry.
- Act as a Designated Responsible Individual (DRI) for networking services and tooling, monitoring production systems, responding to incidents, and driving long-term improvements.
- Stay ahead of emerging trends in AI infrastructure, HPC fabrics (InfiniBand, NVLink, etc.), and software-defined networking to drive innovation.
Krav
- Doctorate Degree in Electrical Engineering, Optical Engineering, Computer Science, Information Technology, or related field; OR Master's Degree with 3+ years technical experience in network design, development, and automation; OR Bachelor's Degree with 4+ years technical experience in network design, development, and automation; OR equivalent experience.
- Ability to meet Microsoft, customer and/or government security screening requirements, including Microsoft Cloud Background Check.
- 6+ years of experience troubleshooting and producing designs that leverage protocols such as OSPF, ISIS, BGP, MPLS, LDP, VRF, etc.
- 4+ years of experience troubleshooting production datacenter Ethernet networks, including switch, port, transceiver, fiber, and physical-layer failures.
- 3+ years of experience troubleshooting high-speed optical connectivity using telemetry such as optical power, BER, FEC, CRC, lane-level status, and link-flap data.
- 2+ years of experience operating or supporting large-scale AI, HPC, or cloud network fabrics using BGP, RoCE, ECMP, or leaf-spine architectures.
- 1+ year(s) experience delivering network designs into production and experience with live site accountability for a network.
- CCNP/CCIE/JNCI certification.
Önskade kvalifikationer
- Doctorate Degree in Electrical Engineering, Optical Engineering, Computer Science, Information Technology, or related field AND 3+ years technical experience in network design, development, and automation; OR Master's Degree with 6+ years technical experience; OR Bachelor's Degree with 8+ years technical experience; OR equivalent experience.
Förmåner
- The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year.
- The base pay range for this role in the San Francisco Bay area and New York City metropolitan area is USD $160,200 - $261,000 per year.
- Certain roles may be eligible for benefits and other compensation.
- Information on additional benefits and pay information can be found at: https://careers.microsoft.com/us/en/us-corporate-pay.
#HPC#AI#Supercomputer#Cloud#Networking#Infrastructure#Distributed Systems#Automation#InfiniBand#RoCE#NVIDIA#AMD GPUs#BGP#MPLS#Data Center