
Machine Learning Engineer, Synthetic Data
Wayve
Job description
Before the detail, here's the challenge you'd help us solve. We build the embodied intelligence that moves real vehicles safely, and the ecosystem a billion machines will run on in the future. Very few people in AI can say this. Every role here, whatever the team, plugs into that. Here’s what this particular role covers. 🛠️ About our Simulation Teams Simulation is advancing end-to-end autonomous driving research. The team’s mission is to accelerate AV2.0 by incubating capabilities that become company-level advantages — generative world models and the synthetic data they produce are one of those. The goal of this role is to build, scale, and optimise next-generation world model architectures (GAIA and successors) and bridge them into high-throughput generation and training infrastructure, so synthetic data can dramatically accelerate autonomy development. You will post-train world models for new embodiments and behaviours (rig transfer, pose transfer, dashcam restaging), generate multimodal synthetic experience at scale, and land that data in the same training stack we use for real driving. You sit between ML research and engineering: collaborating with scientists on architecture and conditioning, and with platform engineers on generation jobs, training artefacts, and how synthetic data is mixed into training. Your work will decide how fast we can train, evaluate, and deploy driving models on vehicles we have barely collected from. 🧠 Your day-to-day - Model work: Post-train or ablate a GAIA rig-transfer or pose-transfer checkpoint (NVS warp, calibration/intrinsics, shortcut/distillation). Inspect failures: black margins, odometry bias, flickering, wrong curvature column. - Generation at scale: Kick off SDS / Flyte jobs for tens of thousands of segments; debug GPU capacity, KV cache, DDIM step count, MCAP/delta-table correctness. - Landing in training: Binarise synthetic into a corpus-compatible table, wire sampling hooks, train RL and read suite + on-road diffs. - Capability expansion: Pose-transfer ELK/AEB sets, dashcam→Alpha3, or a new OEM rig (Nissan Proto 2.1, Gen3, BMW). - Cross-team: Driving-model owners on mix ratios, closed-loop/eval on whether generated MCAPs are usable as suites. - Paper / research alignment: Video generation, novel-view / camera transfer, flow-matching / distillation, data-centric training — only insofar as it changes a checkpoint or a mix we ship. 🧩 What you’ll be working on - Post-train and iterate GAIA-class world models for synthetic-data capabilities: rig transfer (new camera/vehicle embodiments), pose transfer (rewritten ego trajectories), and related conditioning (geometry, calibration, actions). - Own the generation loop: config → large-scale GPU inference → training-ready artefacts, with clear lineage from the model and settings that produced them. - Land synthetic data in driving-model training (behaviour cloning, reward models, RL): binarisation, mix ratios, quality filters, and experiments that measure suite and on-road impact — including when synthetic should replace scarce real rig data. - Diagnose and fix geometry, calibration, and controllability failures (intrinsics/extrinsics, NVS warps, odometry/curvature, flickering, camera-layout artefacts) that determine whether generated video is training-grade. - Improve throughput and yield: inference optimisations (shortcut, distillation, KV cache, step count), valid-generation rate, and self-serve workflows so model developers can request synthetic sets without a specialist. - Expand coverage to new vehicle platforms and safety-critical scenarios (OEM bring-up; Emergency Lane Keeping / Automatic Emergency Braking). - Partner with world-model researchers, infra, and driving-model owners so generation, evaluation, and training stay one system. 🙌 You should apply if - 4+ years in applied ML / research engineering, with a track record of training and shipping neural nets, not only operating data platforms. - Strong Python and PyTorch (or equivalent); comfort with GPU training, debugging, and reading model code. - Hands-on experience with video, generative, or world models (diffusion / flow-matching / autoregressive video, novel-view synthesis, neural rendering, or similar). - Working knowledge of cameras and 3D geometry (multi-camera rigs, intrinsics/extrinsics, warps/reprojection) and why they break generation or downstream training. - Evidence of taking generated or simulated data into a trained downstream model and measuring impact (mix, ablations, failure analysis). - Ability to operate generation or training at real scale (multi-GPU jobs, workflow orchestration, large video artefacts) and to make that path reliable. - Collaborative, experimental working style with researchers and platform engineers; you will own a capability, not a ticket queue. 🌱 Not ticking every box? That’s totally okay! If you’re passionate about autonomy and keen to learn, we encoura
Verified and listed by ActiveJobs. Applications are made directly on Wayve's own career page — we never sit in the middle.