Robots Atlas>ROBOTS ATLAS
Robotics & Hardware

Zero-Copy in ROS 2: an AI Agent Moves a Node to CUDA Buffers

Lady Robot28 September 2026 · 3 min read
Zero-Copy in ROS 2: an AI Agent Moves a Node to CUDA Buffers

NVIDIA published a walkthrough on 22 September for migrating a ROS 2 node to CUDA buffers with the help of an AI agent. The rosidl::Buffer abstraction with a CUDA backend moves data between nodes without serialization and without CPU copies. The worked example is a Depth Anything 3 TensorRT node.

Key takeaways

  • The CUDA backend implements rosidl::Buffer on CUDA Virtual Memory Management
  • Zero-copy requires a shared host, CUDA device, Linux user and supported RMW implementation
  • Outside those conditions the system falls back to the CPU path on its own
  • The agent uses a migrate-node-to-rosidl-buffer skill across six phases
  • The post publishes no performance figures

What rosidl::Buffer actually is

The abstraction represents variable-length primitive arrays, for instance uint8[], in C++ code for ROS 2 Lyrical. The default CPU-backed version behaves like std::vector, so source compatibility is preserved.

Platform vendors can plug in their own backends for externally managed storage. The rosidl_buffer and rosidl_buffer_backend_registry packages live in the ros2/rosidl repository.

When zero-copy actually applies

NVIDIA contributes a CUDA backend that implements rosidl::Buffer storage on CUDA Virtual Memory Management. Zero-copy transport only engages when every condition holds at once.

Publisher prepares the message
Same host, CUDA device, user and RMW
YES
Zero-copy transport
NO
CPU-compatible path
Subscriber receives the data

If any of the four conditions fails — including a supported RMW implementation — the system raises no error. It quietly reverts to the CPU-compatible path. That is why verification is mandatory rather than optional.

What the agent does

The migrate-node-to-rosidl-buffer skill walks the agent through six phases. The worked example is a Depth Anything 3 TensorRT node converting RGB images into floating-point depth predictions.

PhaseWhat happens
1Record the node's starting state
2Confirm message field compatibility and add dependencies
3Trace fields from receipt through to publication
4Audit the data copy boundaries
5Draft a migration plan field by field
6Apply minimal patches that preserve the node interface

How to verify it

C++
// 1. The subscriber declares it will accept GPU-side buffers
rclcpp::SubscriptionOptions opts;
opts.acceptable_buffer_backends = "cuda";
// 2. Output allocated directly in CUDA memory
auto msg = std::make_unique<sensor_msgs::msg::Image>();
cuda_buffer_backend::allocate_buffer(msg->data, output_size);
// 3. Handles for TensorRT — the data never comes down to the host
auto in  = cuda_buffer_backend::from_input_buffer(input->data);
auto out = cuda_buffer_backend::from_output_buffer(msg->data);
// 4. Check on the receiving side
assert(msg->data.get_backend_type() == "cuda");

NVIDIA additionally recommends Nsight Systems to confirm there are no payload-sized host-device transfers at ROS boundaries. The deployment target is the NVIDIA Jetson AGX Thor.

Why it matters

Robotic perception loses most of its time not to compute but to shuttling data between CPU and GPU. Removing those copies at node boundaries buys latency without changing the model or the ROS 2 architecture.

The absence of a benchmark still argues for caution. The post talks about a latency gain but does not publish a single measured value.

What's next

  • The NVIDIA Isaac ROS documentation carries a dedicated section on rosidl::Buffer and buffer backends, including a NITROS migration guide
  • Without published latency and throughput measurements the real gain cannot be assessed
  • The skill generates its own source and sink nodes to exercise both paths: GPU and the CPU fallback

Sources

Share this article