Sr. Systems Design Engineer - Data Center GPU
AMD · MARKHAM, Canada
Job description
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them. AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward. THE ROLE: We are looking for a dynamic, energetic Sr. Systems Design Engineer to join our growing Data Center GPU team. As a key contributor to the success of AMD’s product, you will be part of a leading team to drive and improve AMD’s abilities to deliver the highest quality, industry-leading technologies to market . The Systems Design Engineering team fosters and encourages continuous technical innovation to showcase successes as well as facilitate continuous career development. THE PERSON: In this role, you will drive balanced, scalable, and automated solutions. In this high visibility position, your software systems engineering expertise will be necessary towards Product development, definition, and root cause resolution. KEY RESPONSIBILITIES: p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Executing and contributing to post-silicon validation efforts for SOC-level IP blocks, including test plan development, test execution, coverage tracking, and issue reporting across program milestones p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Debugging hardware and system-level issues found during bring-up, validation, and production phases of SOC programs, with a focus on system IP blocks such as DMA engines, interrupt controllers, and data path logic p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Performing post-silicon debug analysis using scan dump tools and debug reports to investigate hangs, stalls, and error conditions at the IP and system level p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Triaging failures by collecting and correlating debug data across multiple IP domains, with guidance from senior engineers on complex cross-chiplet issues p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Working with multiple teams and tracking test execution to make sure all features are validated and optimized on time p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Working closely with design, firmware, driver, software runtime, and kernel teams to understand IP behavior, error propagation, software-hardware interactions, and power management dependencies p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Developing and maintaining validation test content targeting error handling, fault injection, and power management scenarios p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Engaging in hardware/software modeling and debug frameworks to reproduce and root-cause silicon failures p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Participating in collaborative triage and debug efforts across multiple teams and IP domains, progressively taking ownership of specific IP areas PREFERRED EXPERIENCE: Experience in post-silicon validation, hardware debug, or a related semiconductor engineering role Programming/scripting skills (e.g., C/C++, Python, Perl) p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Debug techniques and methodologies for post-silicon validation p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Experience with board/platform-level debugging, including bring-up, sequencing, analysis, and optimization p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Knowledge of SoC system architecture, including multi-die or chiplet-based designs p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Understanding of cache hierarchies, buffer allocation and management, ring buffers, FIFO structures, credit-based flow control, and tag tracking mechanisms p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Familiarity with DMA, interrupt handling, memory subsystem, or I/O subsystem architectures p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Knowledge of memory management concepts including address translation, virtual memory, TLB operation, and page fault handling p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Familiarity with software runtime environments, kernel-level drivers, or OS-level interfaces that interact with hardware IPs p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Exposure to RAS concepts — error detection, poison propagation, machine check logging, and watchdog mechanisms p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Experience with scan dump analysis or JTAG-based post-silicon debug tools p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Strong analytical/problem-solving skills and pronounced attention to detail p]:inline" style="font-family: arial, helvetica, sans-serif; font-size: 12pt;" data-streamdown="list-item"> Must be a self-starter, able to independently drive tasks to completion and willing to ramp up on new IP domains through documentation and hands-on debug ACADEMIC CREDENTIALS: Bachelors or M asters degree in electrical or computer engineering LOCATION : Markham, ON #LI-SL2 #LI-HYBRID Benefits offered are described: AMD benefits at a glance . AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.
Verified and listed by ActiveJobs. Applications are made directly on AMD's own career page — we never sit in the middle.