
Senior ML Performance Engineer, Inference Optimisation
Wayve
Job description
Before the detail, here's the challenge you'd help us solve. We build the embodied intelligence that moves real vehicles safely, and the ecosystem a billion machines will run on in the future. Very few people in AI can say this. Every role here, whatever the team, plugs into that. Here’s what this particular role covers. 🛠️ About our Engineering Teams Our ML performance team optimises inference for edge accelerators and GPUs, so large transformer-based models run efficiently on low-cost, low-power in-vehicle compute. We work closely with model, platform and deployment teams to turn research models into reliable production systems for Wayve's driving product. 🧠 Your day-to-day This is a hands-on role across ML systems, compilers, runtimes, kernels and embedded deployment. You'll find bottlenecks, implement performance improvements, and work with model developers on performance trade-offs so deployment-aware decisions are made early. 🧩 What you'll be working on - Profiling inference performance across model graphs, compiler/runtime behaviour, kernel execution and memory movement - Implementing and validating optimisations in compilers, runtimes and/or kernels, including fusion, scheduling, quantisation-aware performance and custom kernels - Building benchmarking and regression tests to track performance across models, devices and software releases - Optimising for edge targets such as NVIDIA Orin/Thor and Qualcomm platforms - Contributing to team tooling, documentation and technical discussions around ML performance 🙌 You should apply if - You have experience improving performance in production or production-adjacent systems with latency, memory, bandwidth, power, thermal or cost constraints - You have strong hands-on experience with at least one relevant stack, such as TensorRT, CUDA, Qualcomm QNN, Triton or OpenCL - You are comfortable working from high-level model behaviour down to kernel/runtime-level execution - You have strong software engineering fundamentals, including debugging, profiling, testing and maintainable code - You communicate clearly and work well across ML, systems and deployment teams Nice to have: - Experience deploying or benchmarking ML models on embedded or edge devices - Familiarity with NVIDIA and/or Qualcomm SoCs and performance tooling - Python and C++ proficiency - Experience supporting other engineers or contributing to technical direction within a small team 🌱 Not ticking every box? That’s totally okay! If you’re passionate about autonomy and keen to learn, we encourage you to apply even if you don’t meet every requirement. More about Wayve: 🚀 Wayve is building the leading AI platform for autonomous driving. We are pioneering an end to end AI approach that enables vehicles to learn directly from real world experience, developing the ability to adapt, generalise and improve at scale. Instead of relying on hand coded rules or pre mapped environments, our AI Driver learns to drive by understanding the world around it. The result is technology that navigates complex urban environments with intelligence, precision and natural flow, unlocking meaningful advances in both safety and efficiency. We believe autonomy represents a once in a generation transformation in how people and goods move, comparable to the shift from horses to cars, and from human driven vehicles to intelligent machines. Our ambition is to make autonomy universal. Wayve’s mapless and hardware agnostic AI platform integrates with global OEM partners, enabling continuous software evolution and unlocking advanced levels of automation from L2 plus through to L4 as our core AI model scales. In a race increasingly defined by intelligence and real world learning, Wayve is taking a distinct approach, building a generalisable driving intelligence that can power any vehicle, anywhere. By combining embodied AI with scalable deployment, we are creating technology that can be shaped to each OEM brand and driver experience, accelerating the transition to a safer, more intelligent future of mobility. How we work 💻- Locations & Flexible Working: Our main hubs are in London, Sunnyvale, Yokohama, Herzliya, Vancouver and Leonberg. We operate a hybrid working model that combines in-person collaboration in our dedicated office spaces with focused time working remotely. This gives our teams the connection and energy of working together, alongside the flexibility to do their best work in a way that fits their lives. 🔍 The Interview Process: Our process is clear and respectful of your time: - Initial call / recruiter screen - Deep-dive technical interviews [programming, system design & domain-specific interviews; 3 hours total] - Final interview: mission & values alignment We’ll always explain the format and work around your availability. What’s in it for you (Location dependant): 💰 Salaries benchmarked against the market annually 📈 Meaningful equity, sharing in the
Verified and listed by ActiveJobs. Applications are made directly on Wayve's own career page — we never sit in the middle.