MLOps & Data Engineer
Own the path from raw robot capture to training-ready datasets: triage, versioning, lineage and reproducibility.
About the role
Every robot session generates rich multimodal data, and its value depends entirely on the pipeline behind it. You will own the path from raw capture to training-ready datasets: automated episode triage, versioning and lineage, and the reproducibility discipline that makes training runs trustworthy.
What you'll do
- Build an automated episode labelling layer: filter successes, failures and what needs human review
- Own dataset composition: coverage over task, object and environment space, near-duplicate removal, hard-example mining, balancing
- Run the human-in-the-loop review path
- Own reproducible training runs and regression gates
- Own the consumption side of the data pipeline, downstream of capture
- Perform data augmentation on the obtained datasets
- Work with the Physical AI and Navigation teams so the pipeline matches what they need
What we look for
Some combination of the following:
- You've built a data pipeline for a learned system fed by messy, real-world capture
- Dataset versioning and experiment tracking as daily working practice
- Enough perception literacy to build automated success/failure detection from raw sensor data
- Comfort with timeseries and multimodal robot data at scale
- Judgement about dataset composition, not just volume: you can say why a run improved and point at the data change behind it
Tech stack
- Advanced experience with Python, Linux, Git
- PyTorch, and comfort reading training code you didn't write
- Docker, cloud object storage and infrastructure as code
- LeRobot / RLDS or comparable episode formats
- Dataset versioning and experiment tracking (DVC, MLflow, Weights & Biases or equivalent)
- Timeseries/multimodal tooling
Bonus experience
- Active learning or DAgger-style correction loops
- Training-infrastructure cost optimisation
- Annotation tooling and vision models
