
Director / Fellow – Developer Productivity & Platform, ROCm
AMD · San Jose, California
Job description
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. THE ROLE: ROCm is a 1,200-person engineering organization. The people in it are productive individually. The organization as a whole is not. Developers across firmware, kernel, build, and software teams encounter the same friction points — slow toolchains, fragmented workflows, bespoke local infrastructure assembled team-by-team to fill gaps nobody formally owns — and solve them independently, each time, in different ways. No one has owned this problem horizontally. The problem is real, the scope is rare — a horizontal mandate across a 1,200-person engineering organization, with a clear problem and leadership support to address it. Developer experience roles at this scope don't come up often. Build a function from scratch — you'll establish what developer productivity means at AMD, prove the model, and scale it. Budget exists to grow the team beneath you once the function is validated. This is a founding role, not a slot in an existing org chart. Direct line to leadership — you'll report directly to the head of the Developer Organization, with the access and visibility to move quickly and escalate blockers without bureaucratic overhead. Direct impact on AMD's open-source velocity — ROCm is AMD's bet on the AI software ecosystem. Engineering time is the constraint. This role directly attacks that constraint. Breadth unusual for an IC — firmware, kernel, build, runtime, compiler, ML frameworks — this role touches all of it. If you want depth in one area, this isn't it. If you want horizontal breadth with genuine technical impact across a large, complex organization, it's a rare fit. THE PERSON: We're looking for a senior individual contributor to serve as the horizontal technical lead for developer productivity across the entire ROCm development organization. You'll sit in the Developer Organization (DevOps and QA), reporting directly to the head of that organization, but your mandate spans the full 1,200-person engineering body — surfacing the highest-friction points in the developer experience, and eliminating them: through infrastructure improvements, workflow redesign, or direct fixes to the development stack itself. The goal is measurable, durable improvement in engineering velocity across the company — not for one team, for all of them. This is a solo IC role to start. You'll build the function and prove the model. We have budget to hire supporting engineers beneath you as the work matures, but the first job is to establish what this function owns, what it can change, and what it's worth. KEY RESPONSIBILITIES: Own the developer productivity problem horizontally across ROCm — this is not a team-scoped role. You'll build a continuous, empirical view of where developers across the organization lose time: through direct interviews, instrumentation, friction logging, and analysis of build and CI data. This picture drives your roadmap. Identify and eliminate the highest-leverage friction points — once you can see where the time goes, you prioritize ruthlessly. Some fixes are infrastructure (build caching, environment provisioning, test selection); some are workflow changes (pre-submit gates, LKG baselines, CI reproduction paths); some are direct improvements to the development stack (compiler feedback, incremental build tuning, developer-facing tooling). You'll own and drive all three categories. Build shared tooling that replaces fragmentation — every ROCm component team has assembled workarounds for gaps nobody formally owns. Your job is to surface that work, identify what should become canonical, and drive consolidation into shared, maintained tooling — eliminating the incentive for any team to build their own infrastructure from scratch. Partner with DevOps, QA, and component team leads — you will influence without authority across autonomous teams with their own roadmaps and priorities. Earning credibility requires deep technical fluency alongside organizational effectiveness; this role requires both. Define and instrument the metrics that make productivity legible — build time, incremental compile time, time-to-first-green, environment setup time, onboarding completion rates. If we can't measure it, we can't improve it. You'll own the measurement framework and use it to demonstrate impact to engineering leadership. Translate developer pain into infrastructure investments — you'll be the voice of developer experience on the DevOps roadmap. You'll bring structured evidence of friction points back to the teams who can address them, and help prioritize what gets built next. Build and eventually lead a team — this role starts solo, but the expectation is that you'll prove the model and grow the function. Budget exists to hire supporting engineers beneath you; how you structure that team and when you pull that trigger is part of what you'll own. Developer friction audit — a systematic, org-wide map of where the 1,200 developers in ROCm lose time today: slow builds, environment setup overhead, broken trunk, manual workarounds for missing automation. The output is a prioritized investment list, not a document that sits in a folder. Build acceleration program — identify the highest-impact levers for reducing build latency across ROCm component repos: distributed build caching (sccache or equivalent), incremental build configuration tuning, faster CI test selection. Measure before and after; publish results. Developer environment standardization — define and drive adoption of a reproducible, pre-configured developer environment standard (Dev Containers or equivalent) that eliminates weeks of onboarding archaeology and gives any engineer a working, fully-tooled environment within 15 minutes. CI failure reproduction as a first-class workflow — today, reproducing a CI failure means finding a matching machine manually. You'll drive the infrastructure change that makes any CI failure interactively reproducible on equivalent hardware, eliminating one of the highest-friction moments in the development cycle. Shadow infra consolidation — surface and evaluate the custom scripts, ad-hoc tooling, and bespoke CI configurations built by component teams to fill platform gaps; identify duplication; drive consolidation into shared, owned tooling with formal support. Developer productivity KPIs — instrument and publish a quarterly developer productivity dashboard used by engineering leadership to set priorities and track improvement over time. PREFERRED EXPERIENCE: Software engineering experience with significant depth in developer tooling, platform engineering, build systems, or DevOps Strong hands-on coding ability — this is an IC role; you will be in the code, not just driving process Demonstrated track record of improving developer experience at scale across organizations of 100+ engineers — measurable outcomes required Proven ability to operate horizontally across autonomous teams, building alignment and driving adoption without organizational authority Deep understanding of CI/CD systems, build infrastructure, and developer workflow tooling — you need to understand the problem space at a technical level to earn credibility with the teams you'll work with Strong communication skills: able to translate developer friction into engineering priorities and make a crisp case for investment to senior leadership Experience in GPU or systems software environments (kernel, firmware, runtime, compiler) — understanding what it actually feels like to develop in this stack is a significant advantage Familiarity with the ROCm stack or adjacent GPU compute ecosystems (HIP, CUDA, ML frameworks) Experience with build caching, distributed build systems, or developer environment tooling at scale Fluency with agentic AI workflows (Cursor, Claude, Copilot) — the scope of a horizontal role across a large, fragmented organization is exactly where these tools compound PREFERRED ACADEMIC CREDENTIALS: Master’s degree or PhD in related discipline preferred LOCATION: San Jose, CA #LI-G11 #LI-HYBRID Benefits offered are described: AMD benefits at a glance . AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.
Verified and listed by ActiveJobs. Applications are made directly on AMD's own career page — we never sit in the middle.