
Software Development Engineer — GPU Fleet Management & AI Infrastructure
AMD · San Jose, California
Job description
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. THE ROLE: AMD is looking for an experienced software engineer to help build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure. Fleet Manager provides scheduling, workload orchestration, hardware health management, interactive development environments, and model-serving capabilities across GPU clusters. You will design and build production systems spanning distributed control planes, Kubernetes, GPU scheduling, inference infrastructure, and developer-facing APIs and tools. You will join a team working at the intersection of systems software, cloud infrastructure, and accelerated computing. Your work will directly influence how engineers and customers train, serve, debug, and operate workloads on current and future AMD GPU platforms. THE PERSON: The ideal candidate is a hands-on systems engineer who enjoys solving complex infrastructure problems and turning them into reliable, easy-to-use products. You have strong technical judgment, can reason about distributed-system failure modes, and are comfortable working across service, cluster, and hardware boundaries. You communicate clearly, collaborate effectively across organizations, and can lead substantial projects from architecture through production deployment. You care deeply about correctness, security, operability, and the experience of both end users and platform operators. KEY RESPONSIBILITIES: Design and develop Fleet Manager’s distributed control-plane services, APIs, schedulers, inference gateway, and command-line tools. Build reliable orchestration for GPU training, inference, custom jobs, and interactive development workloads. Develop scalable scheduling and admission-control capabilities, including priority, fairness, topology-aware placement, quotas, backfilling, and multi-node workload coordination. Implement durable reconciliation, lifecycle management, retries, idempotency, and recovery across PostgreSQL and external execution systems. Integrate Fleet Manager with Kubernetes and technologies such as Kueue, JobSet, container runtimes, storage systems, and observability platforms. Help evolve Fleet Manager into a portable orchestration layer capable of supporting Kubernetes, Slurm, Spur, and future execution environments. Develop GPU health, diagnostics, quarantine, and controlled-remediation capabilities using ROCm and AMD hardware telemetry. Build secure multi-tenant infrastructure with strong authentication, authorization, workload isolation, auditing, rate limiting, and least-privilege defaults. Improve the reliability and performance of AI inference services, including routing, streaming, load shedding, health detection, and usage metering. Define and maintain stable APIs, data models, compatibility contracts, and operational procedures. Diagnose complex failures across distributed services, Kubernetes, networking, storage, GPU runtimes, drivers, and hardware. Develop automated unit, integration, failure-injection, and production-readiness tests. Work with AMD architecture, driver, platform, security, and machine-learning software teams to enable current and future GPU products. Participate in new GPU, system, cluster, and software-stack bring-up. Provide technical leadership through design reviews, code reviews, mentoring, and cross-functional problem solving. PREFERRED EXPERIENCE: Strong systems-software development experience in Rust, C++, Go, or a comparable language. Production Rust experience is highly desirable. Experience designing and operating distributed systems, control planes, schedulers, or cloud infrastructure. Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery. Experience building reliable services using REST, streaming, WebSocket, or gRPC APIs. Experience with Kubernetes internals, controllers, operators, scheduling, resource management, or custom resources. Familiarity with workload scheduling technologies such as Kueue, JobSet, Slurm, or other batch and cluster schedulers. Experience with PostgreSQL-backed services, schema evolution, transactions, leader election, and optimistic concurrency. Experience developing command-line tools and stable, user-focused APIs. Understanding of container security, multi-tenant isolation, authentication, authorization, and secrets management. Experience with production observability, including metrics, structured logging, tracing, alerting, and incident diagnosis. Ability to write high-quality, maintainable code with careful attention to correctness, testing, and operational behavior. Experience with source control, continuous integration, automated testing, profiling, and debugging tools. Demonstrated ability to lead technically challenging projects and collaborate across organizational boundaries. Effective written and verbal communication skills. Experience in one or more of the following areas would be beneficial but is not required: AMD GPU architecture, ROCm, HIP, amd-smi , RCCL, or GPU device plugins Distributed AI training and multi-node collective communication Large-model inference using platforms such as vLLM, SGLang, PyTorch, or similar runtimes GPU topology, capacity management, performance analysis, and hardware diagnostics Kubernetes networking, storage, admission control, and workload isolation OpenAI-compatible inference APIs, request routing, streaming, and rate limiting High-performance shared storage and large model or checkpoint management Bare-metal, virtualized, and cloud GPU infrastructure Production security and threat modeling for multi-tenant compute platforms PREFERRED ACADEMIC CREDENTIALS : Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience . This role is not eligible for visa sponsorship. #LI-G11 #LI-HYBRID Benefits offered are described: AMD benefits at a glance . AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.
Verified and listed by ActiveJobs. Applications are made directly on AMD's own career page — we never sit in the middle.