
Data Center Engineer
AMD · Austin, Texas
Job description
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. THE ROLE: The Member of Technical Staff (MTS) Data Center Engineer is a technical leadership role responsible for deploying, validating, troubleshooting, repairing, and optimizing AMD’s next-generation AI, HPC, and data center infrastructure. The role provides technical support both remotely from an AMD office and onsite within AMD customer data center environments. It combines deep technical expertise with broad systems knowledge across servers, GPUs, networking, firmware, power, cooling, and supporting infrastructure. Working with customers, engineering teams, OEMs, ODMs, and service partners, the MTS Data Center Engineer leads complex technical initiatives, influences engineering decisions, and improves product quality, serviceability, operational performance, and customer outcomes. The role operates with a high degree of independence and requires strong technical judgment, systems-level thinking, and cross-functional leadership. THE PERSON: The ideal candidate is an experienced technical leader who thrives in fast-paced environments and can independently solve complex infrastructure challenges. They are equally comfortable providing remote technical support from an AMD office, performing hands-on work onsite in data centers, leading critical troubleshooting efforts, mentoring technical teams, and advising customers and stakeholders. They communicate clearly, make data-driven decisions, and continuously seek opportunities to improve reliability, efficiency, serviceability, and operational excellence. KEY RESPONSIBILITIES: Provide technical leadership for complex deployments, engineering projects, service initiatives, and operational workstreams. Lead or perform the installation, configuration, validation, maintenance, and repair of data center infrastructure. Diagnose and resolve complex server, GPU, networking, firmware, power, cooling, cabling, and system-management issues using structured troubleshooting and root-cause analysis. Perform advanced rack-, cluster-, and node-level validation following deployments, configuration changes, and hardware repairs. Verify firmware, BIOS, BMC, CPLD, networking, and software configurations against approved specifications and standards. Replace, configure, and validate field-replaceable units in accordance with approved technical and safety procedures. Lead site activities, coordinate technical resources, and manage risks, dependencies, milestones, and deliverables. Serve as a technical escalation point and provide coaching, training, and mentorship to engineers and technicians. Coordinate issue resolution with customers, engineering teams, OEMs, ODMs, logistics providers, and service partners. Communicate technical findings, risks, recommendations, and project status to stakeholders and leadership. Develop and improve troubleshooting guides, work instructions, validation procedures, service documentation, and knowledge-base content. Identify recurring failures and process gaps, and lead corrective actions that improve quality, reliability, serviceability, and customer experience. Lead or support incident response, maintenance activities, critical escalations, and service-restoration efforts. Ensure compliance with safety standards, customer security requirements, change-control processes, site procedures, and data-handling requirements. Minimum Qualifications Bachelor’s or master’s degree in engineering, computer science, or a related technical field. Seven or more years of experience supporting data center, server, networking, HPC, AI, or enterprise infrastructure. Advanced knowledge of server architecture, Linux, networking, firmware, power, cooling, and hardware diagnostics. Proven ability to lead complex troubleshooting and root-cause investigations using logs, telemetry, and diagnostic tools. Experience leading cross-functional technical initiatives, influencing decisions, and mentoring technical personnel. Ability to interpret technical documentation and clearly communicate findings, risks, and recommendations. Preferred Qualifications Experience supporting GPU-based AI, HPC, or large-scale cluster environments. Hands-on experience with Linux, server-management interfaces, and Ethernet or InfiniBand networking. Experience with direct-liquid cooling, high-density racks, power distribution, and cooling systems. Experience with deployment, validation, monitoring, automation tools, or scripting languages such as Python, Bash, or PowerShell. Relevant industry or vendor certifications. Physical and Work Requirements Ability to work safely in active data centers and perform physical tasks associated with installing and servicing equipment. Ability to travel to AMD and customer locations as required. Ability to support scheduled maintenance, on-call rotations, weekends, and critical escalations as needed. This role is not eligible for visa sponsorship. Benefits offered are described: AMD benefits at a glance . AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.
Verified and listed by ActiveJobs. Applications are made directly on AMD's own career page — we never sit in the middle.