ActiveJobs
Microsoft

Principal Software Engineer AI Infrastructure

Microsoft

Full-timeOn-sitePosted 6 October 2026
Apply on Company Site →

Job description

Define and drive the architecture and technical roadmap for platform reliability, observability, and operational health. Design unified health and telemetry capabilities that connect customer impact with application, capacity, dependency, infrastructure, and deployment signals. Develop platform safeguards for overload protection, capacity management, routing integrity, configuration consistency, and automated isolation and recovery. Establish engineering practices for production validation, progressive delivery, regression detection, fault testing, and automated rollback. Advance end-to-end request tracing and diagnostics across distributed services, including routing, retries, failover, and asynchronous operations. Lead cross-team architecture efforts, mentor engineers, and turn production learnings into reusable platform capabilities and measurable reliability improvements. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience. Experience with AI/ML serving platforms, high-performance computing, accelerator-based infrastructure or other compute-intensive distributed systems. Experience with distributed tracing, capacity management, load balancing, admission control, retries, backpressure, and graceful degradation. Proficiency in Azure Monitoring systems in the Azure ecosystem would be a plus. Experience designing, building, and operating large-scale distributed systems, cloud services, or other complex production platforms. Experience with reliability engineering, observability, service health, telemetry, and production incident response. Demonstrated ability to lead complex technical initiatives and drive alignment across engineering teams and organizational boundaries. Strong written and verbal communication skills, including the ability to explain technical strategy and architectural decisions to engineers and senior leaders.

Verified and listed by ActiveJobs. Applications are made directly on Microsoft's own career page — we never sit in the middle.