AI Research Infrastructure Engineer
AMD · Austin, Texas
Job description
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them. AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward. THE ROLE: We are seeking an AI Research Infrastructure Engineer to operate, scale, and continuously improve the shared GPU and HPC compute platform behind our AI, ML, and HPC research. You will own the day-to-day health of our SLURM and GPU clusters and work hands-on with researchers to get demanding workloads—including large-scale multi-GPU and multi-node training—running reliably and efficiently. This is a research-enablement role, not a traditional systems-administration role. It requires research literacy—a working understanding of how modern models are trained and where they bottleneck—so you can partner with researchers as a technical peer and directly accelerate their work. The focus is on enabling and operating research infrastructure, not pursuing an independent research agenda. Your scope spans both internal and externally visible compute clusters across the research organization, including AMD's university program clusters and interfacing with external academic collaborators. You may also help coordinate the contractors and systems administrators supporting the environment. Familiarity with modern agentic engineering workflows—tools such as Claude Code, Codex, or Cursor—is also expected. THE PERSON: The ideal candidate should be passionate about software engineering and possess leadership skills to drive sophisticated issues to resolution. Able to communicate effectively and work optimally with different teams across AMD. KEY RESPONSIBILITIES: Own day-to-day operations of the SLURM-managed GPU and HPC clusters, ensuring high availability, utilization, and performance across a multi-user research environment. Partner directly with researchers as a technical peer to get demanding multi-GPU and multi-node workloads running and optimized. Operate and improve the broader compute platform—shared storage, networking, containers, and monitoring—and build automation and self-service workflows that reduce friction for researchers. Support GPU platform and hardware bring-up: validation, enablement, debugging, and operational readiness. Manage AMD's university program clusters and support external academic collaborators alongside internal research users, spanning both internal and externally visible compute clusters. Lead incident response and root-cause analysis, and build runbooks and preventive practices to improve reliability. Set operational standards and best practices, and help coordinate the contractors and systems administrators supporting the environment. Apply agentic coding and operations workflows to improve velocity across deployment, troubleshooting, documentation, and infrastructure management. Experience: Operations- and infrastructure-minded, with a strong bias toward reliability, automation, usability, and reducing operational overhead for researchers. Strong hands-on Linux systems administration in shared/multi-user, production compute environments. Practical, hands-on experience with the SLURM workload manager—partitions, scheduling policy, accounting, node management, and tuning. Experience operating GPU, HPC, or AI/ML research compute environments. Research literacy: a working understanding of how large models are trained and where they bottleneck, alongside researchers as a technical peer on multi-GPU/multi-node workloads. Experience with Docker/containers and Kubernetes, plus automation and system health monitoring. Familiarity with agentic engineering workflows using tools such as Claude Code, Codex, or Cursor. Preferred: DevOps practices: GitHub Actions, self-hosted runners, CI/CD pipelines, and infrastructure automation / Infrastructure as Code. Container registries and reproducible runtime environments. Shared storage, networking, and high-speed interconnects (InfiniBand, RoCE) in multi-user clusters. AMD GPU platforms and the ROCm/RCCL stack—hardware bring-up, architecture-aware debugging, validation, or performance workflows. Supporting developer- or researcher-facing platforms. ACADEMIC CREDENTIALS: Bachelor’s or master’s degree in computer science, computer engineering, electrical engineering, or equivalent preferred LOCATION: Austin, TX This role is not eligible for visa sponsorship. #LI-MR1 Benefits offered are described: AMD benefits at a glance . AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.
Verified and listed by ActiveJobs. Applications are made directly on AMD's own career page — we never sit in the middle.