Software Engineer II
Technology, Data & Digital · Software & Web Development · Software Engineering · DevOps · Site Reliability Engineering
In short
Responsibilities
- Design, implement, test, and operate components of monitoring, alerting, telemetry, diagnostics, and operational intelligence platforms for large-scale Azure services.
- Develop automation that improves incident detection, enrichment, triage, routing, mitigation, recovery, and operational communications.
- Build dashboards and analytics that provide actionable visibility into service health, reliability, and customer impact.
- Contribute to scalable logging, metrics, tracing, diagnostics, and audit infrastructure.
- Develop deployment pipeline and release engineering capabilities, including automated validation, deployment health monitoring, operational readiness checks, safe rollout automation, and release verification.
- Investigate production issues using service telemetry and diagnostics, and partner with engineers to implement durable fixes.
- Participate in DevOps and livesite operations, contributing to service quality, availability, performance, and operational excellence.
- Use AI-assisted engineering and automation tools to improve development, testing, incident investigation, and operational efficiency.
- Collaborate with engineering, operations, and support teams to improve service supportability and the customer experience.
- Follow engineering best practices for security, privacy, quality, reliability, accessibility, and inclusive product development.
Requirements
- Bachelor's Degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- 3+ years of software engineering experience.
- Experience with one or more modern programming languages such as C#, C++, Java, or Python.
- Experience building, testing, shipping, and supporting production software, cloud services, distributed systems, infrastructure platforms, or operational tooling.
- Experience with telemetry, monitoring, alerting, logging, metrics, tracing, diagnostics, automation, or incident-management systems.
- Understanding of distributed systems fundamentals, cloud platforms, or service operations.
- Strong problem-solving, debugging, troubleshooting, analytical, and communication skills.
- Ability to independently own scoped features from design through implementation, validation, deployment, and production support.
- Effectiveness at collaborating with diverse groups of people, openness to feedback, bias for action, and comfort working through ambiguity.
Desired Qualifications
- Experience with operational intelligence, site reliability engineering (SRE), production operations, or service engineering.
- Experience building deployment pipelines, release orchestration systems, rollout automation, deployment health monitoring, or service validation frameworks.
- Experience with service APIs, data-processing systems, event-driven architectures, or reusable infrastructure components.
- Experience applying AI-assisted engineering tools, agents, or intelligent automation to coding, testing, incident investigation, or operational workflows.
- Experience with search technologies, information retrieval, vector databases, retrieval-augmented generation (RAG), or AI-powered applications.
- Ability to reason about distributed services and troubleshoot across application, compute, network, storage, caching, messaging, and load-balancing layers.
- Experience engaging with customers or support teams to understand scenarios and resolve service issues.
- Commitment to security, privacy, accessibility, responsible AI, and customer-focused engineering.
Benefits
- This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Skills
Every company has a mission. What's ours? To empower every person and every organization to achieve more. We believe technology can and should be a force for good and that meaningful innovation contributes to a brighter world in the future and today. Our culture doesn’t just encourage curiosity; it embraces it. Each day we make progress together by showing up as our authentic selves. We show up with a learn-it-all mentality. We show up cheering on others, knowing their success doesn't diminish our own. We show up every day open to learning our own biases, changing our behavior, and inviting…
Company
MicrosoftJob Posted
1 day ago
Employment Type
Full Time
Work mode
Hybrid
Experience Level
Mid-Senior
Locations
United States
Qualification
Applicants
Be an early applicant
