Ros2probe: observe ROS 2 traffic without the probe effect (swap ros2 -> rp)

ros2probe is a drop-in replacement for rosbag2 and ros2 topic that does not perturb the system it observes. It records and monitors ROS 2 traffic from outside the DDS domain. No extra subscriber, no observer-induced drops, far less CPU and memory. It reports exactly the loss the real subscriber saw, and the CLI mirrors ROS 2, so you just swap ros2 for rp.

Code · Paper (arXiv) · Project page

Why this exists

Every standard observer (rosbag2, ros2 topic echo / hz, the ros2 daemon,
DDS vendor monitors) sees your data by subscribing. That adds a DataReader, the
publisher sends an extra copy, and near link saturation that copy steals
bandwidth from the real subscriber. The observer also reads a different copy, so
the loss it reports is not the loss your subscriber actually experienced. This
is structural to DDS pub/sub, not a bug in any one tool.

What ros2probe does differently

  • Passive eBPF wire tap. Reads a kernel copy of RTPS off the wire.
  • Never joins the domain. No participant, no DataReader, nothing added to the wire.
  • Reads the same packets the subscriber gets. The loss and latency it reports match what the subscriber actually saw (recall 1.0).
  • Full reconstruction in userspace. Topic graph, per-topic metrics, and message streams, independent of the DDS vendor.

Results on real hardware

3 platforms (laptop, Jetson, Raspberry Pi), 2 DDS implementations (Fast DDS, Cyclone DDS), 7 workloads, wired and wireless, two QoS settings.

ros2probe existing ros2 tools
Loss it causes on the subscriber (recording the full near-GbE workload) 0% up to 75.5% (rosbag2)
Loss it reports vs what the subscriber saw exact, recall 1.0 0.09 (rosbag2 at 10% loss)
Discovery graph perturbation within 0.5% up to 2.6x inflation
Observer CPU up to 7x lower baseline
Observer memory up to 28x lower (1.7 MB vs ~47 MB) baseline

On a Raspberry Pi 4B, ros2 topic hz saturates a CPU core while ros2probe stays under 30%.

Usage. You already know the commands

Install once, then replace ros2 with rp:

ros2 topic hz   /scan    ->    rp topic hz   /scan
ros2 bag record /scan    ->    rp bag record /scan

Recordings are written as MCAP and replay with ros2 bag play. Scripts and CI
that parse hz output keep working unchanged.

ros2probe also ships a GUI (rp gui) with a live ROS graph, a per-topic monitor, and an MCAP recorder.

Status

  • Works today on RTPS-based DDS. Tested on Fast DDS and Cyclone DDS.
  • Shared-memory (SHM) transport is a structural limit, not a blind spot. The
    topic graph is still recovered passively from discovery on the network, but
    SHM payloads never reach the wire, so observing them falls back to a
    short-lived, namespace-isolated shadow subscriber that joins the domain. For
    those topics ros2probe still works, but gives up its probe-effect-free
    guarantee. Network-transported topics keep the full benefit.
  • Zenoh (rmw_zenoh) uses a different wire protocol and is planned. RViz
    integration is on the roadmap.

Links

Feedback, issues, and real use cases are very welcome.
If you are curious about my other work, see https://hun0130.github.io/.

15 Likes

That’s such a great idea!

FYI, @pablothepenguin , seems relevant to your ros_tap tool.

nice work, the shadow subscriber trick is really clever!

1 thing worth fixing though: the README says the filter attaches to non-loopback interfaces only and that SHM-only topics are excluded from recording, but from what I see the code always captures on lo and spawns the shadow sub for SHM topics, so single-host actually works better than the README suggests, right? :smiley:

I’d like to try the passive hz/delay metrics as an input for health monitoring in ros2_medkit, where single-host is the common case

@sanghoon_lee I will report back once we’ve given it a proper test :sweat_smile:

Thanks for the kind words, and great catch on the README. You’re right, the code does capture on lo and spawns the shadow sub for SHM topics, so single-host works better than the docs claim. I’ll fix it to match the actual behavior.

The ros2_medkit idea sounds really interesting. Using the passive hz/delay metrics as a health-monitoring input is exactly the kind of use case I was hoping people would find. Please do give it a spin and let me know how it goes, and feel free to throw any rough edges or feature requests my way. Looking forward to your report!

1 Like

Thanks, that means a lot! And thanks for the pointer to ros_tap. I went and read through it, and I think they’re actually after fairly different things.

From what I can tell, ros_tap joins the DDS network as a CycloneDDS participant and subscribes to stream telemetry (JSONL to stdout, disk, or S3), which is a clean fit for zero-config fleet capture from any machine. ros2probe goes the opposite direction and sits below the middleware, reading a kernel copy of the RTPS packets via eBPF, so it never joins the graph at all. That no-participant part is really the whole point for us, since adding a subscriber is exactly the probe effect we’re trying to avoid.

So different layer and different goal, but I appreciate the connection.

Yes, I know. And I think that’s the better approach for the recording ros_tap aims to do as well, so thought I Pablo might be interested.

This is excellent work. The “observer effect” problem in ROS 2 tooling is very real, especially near bandwidth or CPU saturation. A non-intrusive RTPS/eBPF-level recorder is a very valuable layer.

One thought: this seems highly complementary to a different class of runtime assurance tools that operate above transport-level observability.ros2probe answers questions like:

Did the real subscriber receive the packet?
Was there DDS/RTPS loss?
Did the observer perturb the graph or add load?
What latency/loss did the subscriber actually see?

There is another failure plane where the transport can be perfectly healthy, but the control intent is no longer semantically or physically healthy.

For example, in PX4 Offboard autonomy, the network may deliver every /fmu/in/trajectory_setpoint packet correctly, but the upstream planner / VIO / perception stack may be delayed, bursty, or re-emitting setpoints that are stale relative to the current vehicle state. In that case, a transport-level probe can correctly report “delivery is fine,” while the autonomy stack still needs a boundary-level check:

Is the setpoint stream fresh?
Is it jittered?
Is it consistent with the current Offboard mode?
Is the vehicle response physically matching the intent stream?

I have been experimenting with this complementary layer in a small ROS 2 / PX4 project called AFIO, currently reframing it as Autonomy Flight Integrity Observer. It is a passive Offboard boundary observer that watches:

/fmu/in/trajectory_setpoint
/fmu/in/offboard_control_mode
/fmu/out/vehicle_odometry

and publishes standard /diagnostics plus CSV labels such as:

setpointAgeMs
setpointJitterMs
staleStreams
positionTrackingResidual
velocityTrackingResidual
flightResidual
dominantCause

In controlled PX4/Gazebo latency-injection tests, the transport can remain syntactically valid while the Offboard boundary transitions from healthy → SETPOINT_JITTER → STALE_STREAM.

I see a very natural integration path:

ros2probe:
  non-intrusive capture / MCAP / true subscriber-side transport metrics

AFIO:
  domain-level residual analysis on trajectory_setpoint + odometry + mode semantics

That combination could give both:

Did the subscriber receive the data?
and
Was the received data still a valid control intent for the vehicle?

Question: does ros2probe expose reconstructed message payloads and timestamps in a way that downstream tools can consume live or from MCAP? If so, it would be very interesting to run Offboard boundary-integrity analysis on top of ros2probe recordings without adding any extra ROS 2 subscribers.

AFIO repo for reference: https://github.com/ZC502/ai_flight_integrity_observer.git

Sounds very promising, any chance that this could get integrated into the official repos? @emersonknapp @mjcarroll

Rosbag under ros2 is really an issue if you have many topics.

Hi everyone — a quick ros2probe update!

Zenoh support has landed in v0.2.0. ros2probe can now observe rmw_zenoh traffic over both TCP and UDP, alongside Fast DDS and Cyclone DDS.

More details: ros2probe · Non-intrusive Observability for ROS 2

3 Likes

@zc_Liu, I’m very sorry for the delayed reply. I only just saw your thoughtful message.

If the project is still ongoing, could you please send more details about the integration you have in mind to leesh2913@dgist.ac.kr? Collaboration is always welcome!

@lars, if ros2probe could be integrated into the official repositories, it would be a tremendous honor for us!
We are continuously maintaining and expanding the project. We are currently working on an RViz extension and a layer-by-layer latency measurement tool as well.

@sanghoon_lee Thanks, and congratulations on the v0.2.0 release !​
The integration idea I had in mind was fairly simple: use ros2probe’s passive RTPS/Zenoh message reconstruction as the data source for higher-level boundary-integrity checks, so metrics such as setpoint freshness/jitter and command-response residuals could be computed without adding another ROS 2 subscriber.​
The project has since evolved from AFIO into a broader OBIO/runtime-assurance direction. But my current development bandwidth is focused on LLM inference and deployment diagnostics, so I probably can’t drive a deep ros2probe integration in the near term.​
But I still think the two layers are naturally complementary, and I’d be happy to revisit it if a concrete autonomy/Offboard use case comes up. Thanks again for reaching out!

@zc_Liu, thanks for laying that out so carefully, and no pressure at all about
the bandwidth on your side.

I think you are pointing at something real, and it is close to what our group
has been working on from the other direction. Your framing is that the transport
can be syntactically valid while the control intent is no longer physically
valid, and that the check belongs at the Offboard boundary rather than inside
the transport. We agree, and we have been pushing on the step right after it.
Once a boundary observer can tell that the setpoint stream is stale or jittered,
the question becomes what the system does with that verdict. Detection on its
own still leaves the stale command on its way to the vehicle.

We wrote that argument up here, in case it is useful reading:

Harness Engineering for Physical AI: Robot Middleware Is the Harness Layer

The short version is that robot middleware is the natural host for the
enforcement, because it is the lowest layer that abstracts control, computing
and communication at the same time. We split the enforcement into three
functions. Isolation bounds execution and transmission, which is where your
setpointAgeMs and setpointJitterMs live. Projection gates the output at the
moment of emission, so a command that fails the contract is never published.
Transfer falls back to a verified baseline when a check fails. Your
healthy → SETPOINT_JITTER → STALE_STREAM progression is exactly the kind of
state sequence such a contract is meant to act on, rather than only label.

We are currently running a case study on a ROS 2 navigation stack, binding
output, timing and operating conditions into a single profile and measuring what
happens when a value-valid but late command arrives, including how the same
profile behaves across different policies and different DDS backends. The
implementation, the profiles and the reproduction scripts are planned for public
release this year.

Which leaves three layers stacking up in this thread, and that is a good place
to be:

  • ros2probe — did the data actually arrive, and what did observing it cost
  • your OBIO direction — was what arrived still a valid control intent
  • the harness work — and what should the system do about it before the
    command reaches the actuator

If you ever want to pick the integration back up, or just compare notes on how
you are deciding dominantCause, the door is open. leesh2913@dgist.ac.kr

Thanks — I read the paper and enjoyed the framing.
One thing I kept running into while working on OBIO is that sequence-level drift can emerge even when individual commands or messages still look valid, while the control/runtime stack often has limited visibility into that evolving behavior.
That observation is actually one reason I’ve recently shifted part of my attention to LLM inference, where the same broader question appears in a very different form: whether a deployed low-precision system preserves the sequential behavior of its reference.
I hope some of what comes out of that work may eventually be useful to your harness direction as well.
My active bandwidth is there for now, but if your project is still moving forward when I return to the robotics side, I’d be happy to reconnect — especially around how an observer-side verdict such as dominantCause could feed a middleware enforcement contract.
In the meantime, please feel free to borrow anything useful from the OBIO boundary-observer approach. Best of luck with the project!