ActiveJobs
Meta

Cloud Production Platform Engineer

Meta · Los Lunas, NM

Full-timeOn-sitePosted 6 October 2026
Apply on Company Site →

Job description

Meta is seeking a experienced hardware engineering professional with technical leadership experience and skills in Server Hardware, high performance Compute, Storage, Accelerators (GPU), and/or Networking, ideally in a data center environment. Our data centers, the global fleet of servers installed in them, and the growing cloud capacity Meta operates alongside them are the foundation upon which our rapidly scaling infrastructure efficiently operates and on which our innovative services are delivered. Our Production Platform Engineers are responsible for the performance of the compute, storage, and accelerator (GPU) platforms that run Meta's production workloads, both in our own data centers and on capacity Meta consumes from cloud providers. They are responsible for the health of the server fleet, from NPI production operational test through end of life. Responsibilities include identifying systemic hardware, firmware, and tooling issues; engaging in hands on problem solving; and collaborating effectively with engineering, tooling, and onsite provisioning and break fix teams to improve performance of the fleet. This position carries the cloud platform scope for the team. Meta's cloud footprint is growing toward the scale of the fleet we run ourselves, across several cloud providers, and product groups expect the same reliability and availability from it. Because Meta does not design this hardware and does not have physical access to it, the engineer in this role works through each cloud provider's own repair and escalation process, reaches diagnostic conclusions with partial telemetry, and builds the technical evidence that supports cloud provider accountability. This remains a platform engineering role across both environments. It is not a cloud only role. This position requires a high degree of understanding of our platforms, leveraging technical expertise and using data to identify trends across the fleet and proactively identify and mitigate performance and reliability issues. Experience organizing and prioritizing work across multiple teams in a large-scale, distributed environment and experience partnering with cross-functional teams to drive outcomes in a large-scale, distributed environment.

Verified and listed by ActiveJobs. Applications are made directly on Meta's own career page — we never sit in the middle.