
Senior AI Performance Architect United States California Mountain View
Microsoft
Job description
Analyze frontier training and inference workloads to identify key performance requirements and bottlenecks. Develop and improve analytical models, simulators, and data-analysis tools for AI system performance. Evaluate trade-offs in performance, scalability, utilization, capacity, and cost across accelerator, memory, network, and software configurations. Study hardware-software interactions in large-scale AI systems and translate workload insights into architecture requirements. Work with MSI and partner teams to evaluate hardware-software co-design options for training and inference. Develop and run targeted microbenchmarks on silicon to measure compute, memory, communication, kernel, and synchronization performance. Compare analytical model predictions with silicon and end-to-end workload measurements, identify gaps, and improve model accuracy. Communicate findings and recommendations through clear data visualizations, technical reports, and architecture decision materials. Master's Degree in Computer Science, Electrical Engineering, Computer Engineering, Applied Mathematics, or a related field AND 3+ years of technical engineering experience; OR Bachelor's Degree in a related field AND 5+ years of technical engineering experience; OR equivalent experience. Doctorate in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 3+ years technical engineering experience OR Master's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 6+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Mechanical Engineering, or related field AND 8+ years technical engineering experience OR equivalent experience Experience in AI systems performance, computer architecture, performance modeling, high-performance computing, distributed systems, or accelerator software. Strong understanding of computer architecture and the performance interactions among compute, memory, communication, and software. Experience with analytical performance modeling, simulation, profiling, workload characterization, or system performance analysis. Hands-on programming experience in Python and at least one systems or accelerator programming environment such as C++, CUDA, Triton, ROCm/HIP, or an equivalent platform. Experience analyzing AI training or inference workloads, including latency, throughput, capacity, utilization, and scaling behavior. Experience with transformer-based models, large language model training and inference, attention, KV-cache behavior, mixture-of-experts, or speculative decoding. Experience modeling or optimizing distributed execution strategies such as tensor, pipeline, expert, or data parallelism, or prefill/decode disaggregation. Knowledge of accelerator architecture, HBM and memory hierarchies, scale-up and scale-out networks, PCIe or other high-speed I/O, kernels, compilers, and runtimes. Experience correlating analytical or simulation models with silicon measurements and performing root-cause analysis of performance gaps. Experience with data analysis and visualization for performance studies and architecture decisions.
Verified and listed by ActiveJobs. Applications are made directly on Microsoft's own career page — we never sit in the middle.