ActiveJobs

Principal Systems Validation Engineer – Data Center GPU System Stress

AMD · VANCOUVER, Canada

Full-timeHybridPosted 3 September 2026
Apply on Company Site →

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. THE ROLE Our Data Center GPU (DCGPU) Validation and Engineering team ensures the quality, reliability, security, and performance of AMD's industry-leading AI, ML, and HPC products. Working closely with architecture, design, product engineering, software, and hardware teams, we develop the validation methodologies, automation frameworks, and system-level debug capabilities that enable successful silicon bring-up, product readiness, and exceptional customer quality. Our products demand rigorous testing for optimum performance and reliability under maximum stress. You will play a key role in shaping those strategies and their implementation. THE PERSON We are seeking a Principal System Validation Engineer to help drive product-level system and stress testing across AMD's next-generation platforms. In this role you will define and drive the validation strategy for system stability and stress at scale, translate that strategy into test content and automation, and drive collaboration across engineering organizations. You will operate with extensive silicon-to-systems understanding, driving complex system-level integration challenges and ensuring stable system configuration under maximum stress workloads. The Systems Design Engineering team fosters and encourages continuous technical innovation to showcase successes as well as facilitate continuous career development. KEY RESPONSIBILITIES Define, drive, and evolve AMD's system stability and stress validation strategy, and drive its implementation across test content, coverage models, and automation frameworks to ensure test coverage and configurations lead to full system stability. Develop a deep understanding of our silicon SoCs, board-level designs, system-level designs/interfaces, and firmware/software stacks to drive the development of system stress tests and complex issue investigations. Work across multiple validation, firmware/software, silicon design, and verification teams to develop and execute robust validation test plans at the stress and scale levels that meet our customer requirements. Provide technical leadership for complex system-level challenges spanning HW/FW/SW interactions and apply those learnings back into stronger test coverage, improved tools, and optimized workflows. Drive technical innovation across validation, including design and development of AI-assisted validation workflows and tooling, at-scale validation environments, and data-driven quality methodologies that improve debug efficiency, execution speed, and product readiness. PREFERRED EXPERIENCE Deep understanding of modern GPU, SoC, and server platform architectures, including experience with post-silicon bring-up, silicon to system-level validation, and AI/ML accelerator or large-scale data center platforms. Extensive knowledge of system validation strategy, with a proven ability to develop validation architectures, methodologies, test strategies, automation frameworks, and end- to-end infrastructure for complex data center and semiconductor product validation Extensive experience with SoC/board/platform-level debug — including delivery, sequencing, analysis, and optimization — with a structured approach to debug workflows. Strong analytical/problem-solving skills, pronounced attention to detail, and a passion for innovation and continuous improvement. History of process improvements and driving early critical-coverage enablement. Hands on programming/scripting/debug skills (e.g., C/C++, Perl, Ruby, Python), along with working knowledge of Linux and Windows server environments. ACADEMIC CREDENTIALS: Bachelor’s or Master’s Degree in Computer Engineering or Electrical Engineering LOCATION: Vancouver, BC #LI-EV1 #LI-HYBRID Benefits offered are described: AMD benefits at a glance . AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.

Verified and listed by ActiveJobs. Applications are made directly on AMD's own career page — we never sit in the middle.