
Principal Engineer, Model Development Platform
Wayve
Job description
Before the detail, here's the challenge you'd help us solve. We build the embodied intelligence that moves real vehicles safely, and the ecosystem a billion machines will run on in the future. Very few people in AI can say this. Every role here, whatever the team, plugs into that. Here’s what this particular role covers. 🛠️ About our Engineering Teams The Model Development Platform team builds the infrastructure and tooling behind Wayve's AI model lifecycle, from data ingestion and training to experiment scheduling and on-road testing. Our work spans AI research, large-scale distributed systems and robotic operations, and lets researchers and engineers iterate fast and deploy autonomous driving models safely. 🧠 Your day-to-day As Principal Engineer, you'll own the end-to-end architecture of the platform and keep it reliable, scalable and coherent. You'll partner with the Head of Model Dev Platform to set and execute the technical vision, aligning infrastructure and tooling with company goals. You'll lead by example, going deep across web applications, distributed compute, ML Ops, data pipelines and optimisation algorithms, and through architecture and mentorship you'll help teams build platform capabilities that measurably speed up model development and fleet learning. 🧩 What you'll be working on - System architecture and reliability: designing and evolving the platform's architecture for reliability, observability and scalability, setting performance, latency and availability targets, and driving the engineering standards to meet them - Cross-domain technical leadership: unifying the platform across front-end UIs, distributed training, Spark data pipelines and optimisation-based experiment scheduling, so systems work together cleanly - Hands-on problem solving: taking on the hardest problems across subteams, leading architectural reviews and proposing pragmatic solutions that balance innovation with operational simplicity - Experimentation and scheduling systems: building systems that optimise how models are tested in simulation and on-road, using techniques like linear programming and heuristic optimisation to balance hardware, safety and research priorities while improving throughput and turnaround - Data and compute infrastructure: architecting pipelines that ingest, transform and enrich petabytes of fleet sensor data, and driving efficient compute use across GPU, CPU, cloud and edge for prototyping and large-scale training - Strategic collaboration: working with Product, Research and Operations to align architecture with user needs, and co-owning the platform's long-term roadmap 🙌 You should apply if - You have 10+ years designing and building large-scale distributed systems, ML/AI infrastructure, full-stack web applications or developer platforms, including at least 3 years as a staff or principal-level engineer - You have designed systems spanning web platforms, ML pipelines and large-scale compute orchestration (e.g. Spark, Ray, Kubernetes, Airflow, MLflow) - You have driven platform reliability improvements, defined SLAs/SLOs, and built self-healing, observable systems that run at "four nines" availability or better - You understand distributed computing, workflow orchestration, data modelling and API design in depth, and can write and review production-quality code - You communicate well across functions and can guide engineers, managers and researchers toward a unified technical direction - You have mentored engineers across levels and built a culture of engineering excellence Nice to have: - Experience applying algorithmic or mathematical optimisation (e.g. linear programming, graph algorithms) to operational or scheduling problems - Familiarity with end-to-end model lifecycle tooling, from data ingestion and training CI to model artifact tracking and evaluation workflows - Prior exposure to autonomous systems, robotics or other safety-critical domains - Experience with modern web frameworks (e.g. React, Flask, FastAPI) and how they integrate with backend systems - Understanding of data privacy, compliance and secure handling practices for large-scale sensor data 🌱 Not ticking every box? That’s totally okay! If you’re passionate about autonomy and keen to learn, we encourage you to apply even if you don’t meet every requirement. More about Wayve: 🚀 Wayve is building the leading AI platform for autonomous driving. We are pioneering an end to end AI approach that enables vehicles to learn directly from real world experience, developing the ability to adapt, generalise and improve at scale. Instead of relying on hand coded rules or pre mapped environments, our AI Driver learns to drive by understanding the world around it. The result is technology that navigates complex urban environments with intelligence, precision and natural flow, unlocking meaningful advances in both safety and efficiency. We believe autonomy represents a once in a genera
Verified and listed by ActiveJobs. Applications are made directly on Wayve's own career page — we never sit in the middle.