ROS 2 Task Resilience — experimental mission persistence and restart recovery with Nav2

Hi everyone,

I’ve published an experimental developer preview of ROS 2 Task Resilience, a personal open-source project for durable mission state and recovery after a supervisor process restarts.

The problem I’m exploring is what happens to unfinished work when the process managing a mission disappears. On restart, the application needs to distinguish completed steps, actions that were already dispatched, and outcomes it cannot establish, without simply starting the mission over or sending the same goal again.

The current hardware-free Nav2 demo runs a two-step mission: visit checkpoint A, then return to base. In the restart scenario, checkpoint A is already complete, and the supervisor exits at a controlled checkpoint after submitting the return goal but before recording its result. Nav2 stays running. A separate supervisor process then opens the same database and recovers the retained result for that dispatch without resending the goal.

The implementation includes a ROS-independent SQLite core for mission and step state, dispatch identity, and execution-attempt accounting. A separate ROS-facing package provides single-host execution ownership, a thin NavigateToPose adapter, and a small run/resume/status CLI. It does not replace Nav2’s planning or navigation behavior.

There is also a fake-action-server quickstart that demonstrates separate-process result recovery without launching the Nav2 stack. It still needs ROS and the relevant action message packages.

This is an early preview, tested on ROS 2 Jazzy with Ubuntu 24.04 under WSL2, using simulation only. The preview supervisor supports one mission per dedicated database. Cancellation integration and automatic resubmission/retries are not implemented, and there is no exactly-once guarantee. Stopping the supervisor does not cancel remote navigation.

Recovery depends on evidence still being available from the action server. When it is not, the invocation can remain UNRESOLVED rather than assume success or resend an uncertain goal.

I’d appreciate feedback on whether this addresses restart problems you have encountered, which existing tools or approaches I should compare against, and whether the quickstart runs as documented. Feedback on the API and CLI shape would also be helpful.

Thanks,
Dennis