
Principal Software Engineer Aks Flex High Scale United States Multiple Locations Multiple Locations
Microsoft
Job description
Own the architecture for control plane scalability and the flex nodes and Unbounded platform, and help define the multi-year technical roadmap for AKS Flex in partnership with product and engineering leadership. Lead cross-organization design efforts spanning AKS, Azure compute, storage, networking, and identity, driving alignment on interfaces, tradeoffs, and sequencing. Make and defend key architectural decisions around correctness, durability, security, scale limits, and failure modes, and establish the design review and quality bar for the area. Mentor and grow senior engineers, and raise the technical capability of the broader team. Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, Go, Rust, C++, C#, Java, or Python OR equivalent experience Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, Go, Rust, C++, C#, Java, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, Go, Rust, C++, C#, Java, or Python OR equivalent experience. Deep understanding of etcd internals, including Raft consensus, MVCC, the watch mechanism, compaction and defragmentation, and the bbolt storage backend, as well as experience operating etcd at scale. Deep understanding of storage systems and databases, including storage engines (B-trees, LSM trees), write-ahead logging, replication, consistency models, transactions, and performance tuning. Experience with Kubernetes control plane internals: API server, watch cache, controller patterns, API priority and fairness, and scalability limits. Experience with performance engineering: profiling, load testing, capacity modeling, and diagnosing latency and throughput issues in production. Analyze and optimize control plane bottlenecks such as write amplification, compaction and defragmentation, watch fan-out, list/relist storms, request prioritization, and storage and memory pressure. Build scale and performance testing infrastructure, benchmarks, and SLOs that let us measure control plane limits and prevent regressions before they reach customers. Build automation for safe etcd operations at fleet scale: backup and restore, member replacement, quorum recovery, upgrades, and repair. Design and build the systems that let worker nodes in other Azure regions, on-premises data centers, edge locations, and other clouds securely join an AKS cluster and behave as standard node pools. Build node provisioning and day-2 lifecycle automation across diverse substrates: SSH-based bootstrap, cloud API provisioning in response to unschedulable pods, and bare-metal PXE boot with BMC power management and TPM-based attestation. Design Kubernetes-native APIs, controllers, and CRDs that make remote machines declaratively manageable with standard tooling, and that scale efficiently as fleets grow. Contribute to and help lead the Unbounded open-source project, including design discussions, code reviews, and community engagement. Participate in production operations and on-call rotations, turning operational learnings into durable engineering improvements.
Verified and listed by ActiveJobs. Applications are made directly on Microsoft's own career page — we never sit in the middle.