Mission Orchestration Engineer
Own the loop between autonomy and human pilots: mission state machine, stuck detection and system contracts.
About the role
Our robots combine autonomy with expert human pilots, and the loop between them is a product in itself: dispatch, run, detect when stuck, call for help, return, stand by. You will own that loop, the mission state machine behind it, and the software interface connecting pilots to robots. This is deep integration engineering: making a multi-team robot stack compose into one coherent system.
What you'll do
- Own and extend the robot state machine: stand, walk, autonomous behaviour, human take-over
- Build robust stuck detection: what counts as stuck, how quickly it's noticed, and what the robot does while waiting
- Turn timeouts, watchdogs and escalation paths into one consistent policy instead of per-module ad hoc logic
- Define the contracts between autonomy, controls, perception and pilot tooling
- Define system-level health metrics that describe the whole loop, and use them to decide where to invest next
- Work daily with Controls, Autonomy and Operator Experience to keep their interfaces coherent
What we look for
Some combination of the following:
- Senior experience building distributed systems on real hardware
- Explicit state machines you have designed, shipped and debugged in the field
- Comfort owning seams between teams, and turning competing needs into interfaces that hold up
- Judgment about failure handling: retry, degrade, or hand to a human — and make the choice legible afterwards
- Production C++ and Python, plus enough web literacy to follow the loop to the pilot's screen
- Observability instincts, and a willingness to spend time on the floor with the robots
Tech stack
- Advanced experience with Python, C++, Linux, Git
- Docker, CI/CD, Kubernetes
- DDS communication systems (ROS2, CycloneDDS, etc.)
- Logging and replay of full missions, so a stuck event can be reconstructed after the fact
- Metrics and dashboards for loop-level health
Bonus experience
- Teleoperation products: latency budgets, handover UX
- Fleet operations at scale, including on-call and incident review
- Fault management from aerospace, automotive or industrial automation
- Legged or mobile manipulation platforms and their recovery behaviours
- Simulation tools like NVIDIA Isaac Sim or MuJoCo
