October 4, 2026
Robotics Governance

ROS 2 in Production: Migration, Real-Time, and DDS Tuning

ROS 2 in Production Migration, Real-Time, and DDS Tuning

ROS 2 becomes production-ready when teams treat DDS tuning, QoS selection, and real-time kernel configuration as first-class engineering work rather than afterthoughts. This guide walks through a practical migration path from ROS 1, the specific kernel and middleware settings that cut latency and dropped messages, and the pitfalls that most often derail a first production rollout.
Migration stageROS 1 approachROS 2 production approach
Communication backboneSingle master, TCPROS/UDPROSDecentralized DDS, no single point of failure
Build toolingcatkin_make / catkin buildcolcon with ament build type
Message delivery guaranteeImplicit, connection-basedExplicit QoS policies per topic
Real-time supportBolted on via external patchesNative support through rmw and PREEMPT_RT kernels
Lifecycle managementManual, ad hocManaged nodes with defined states

Why ROS 2 Production Deployments Are Different From Prototypes

A ROS 2 prototype running on a developer laptop and a ROS 2 stack running a fleet of warehouse robots or an outdoor inspection platform are, technically, the same framework. Practically, they are almost different products. The prototype tolerates dropped messages, unbounded queues, and a middleware vendor’s default settings. The production system cannot. It has to survive network congestion, multi-hour uptime without memory growth, degraded Wi-Fi near metal shelving, and a fleet manager pushing software updates while robots are mid-task.

The good news is that ROS 2 was designed with this gap in mind. It replaced the single-master architecture of ROS 1 with the Data Distribution Service, a decentralized publish-subscribe standard that removes the central point of failure and exposes granular Quality of Service controls per topic. The catch is that none of that capability helps unless someone actually configures it. Out of the box, ROS 2 defaults are tuned for developer convenience, not for a robot operating unattended in a warehouse at 2 a.m.

This article is a working guide for the three areas that make or break a production ROS 2 deployment: migrating cleanly from ROS 1 when that applies, configuring the operating system and middleware for real-time behavior, and tuning DDS and QoS settings so that safety-relevant and best-effort traffic are treated appropriately differently.

Step 1: Plan the ROS 1 to ROS 2 Migration as an Audit, Not a Port

Teams that treat migration as a line-by-line port of ROS 1 nodes into ROS 2 syntax tend to under-budget the effort by a wide margin. The more reliable approach starts with an audit.

1.1 Inventory Every Node, Message Type, and Dependency

Before converting a single line of code, list every node in the system, every custom message and service definition, and every third-party package the stack depends on. ROS 2 does not support the old package.xml format 1 specification, so manifests need to be updated to format 2 or later regardless of anything else. If a dependency has no ROS 2 release, that is a blocking issue to resolve early, not something to discover mid-migration.

1.2 Convert the Build System First

ROS 2 uses colcon instead of catkin_make or catkin build, and it expects a newer minimum CMake version than most ROS 1 workspaces use. Get a package building cleanly under colcon and ament before touching runtime logic. This isolates build-system errors from logic errors and makes debugging much faster.

1.3 Migrate Incrementally With a Bridge

Very few teams can freeze feature development for a full-stack rewrite. The ros1_bridge package lets ROS 1 and ROS 2 nodes exchange messages during a transition period, so subsystems can be migrated one at a time while the rest of the fleet keeps running on the existing stack. Start with leaf nodes that have few dependents — sensor drivers and simple utility nodes — before touching the coordination and planning layers that everything else depends on.

1.4 Re-verify Timing Assumptions

ROS 1 code frequently made implicit assumptions about message ordering and delivery that were true only because of how the master and TCPROS behaved. Those assumptions do not automatically carry over. Any node that assumes “the last message published is always the first one received” needs to be re-checked against ROS 2’s QoS-driven delivery model.

Common mistake

Teams frequently migrate node logic first and treat build tooling, QoS, and dependency audits as cleanup work to handle “later.” This backfires because build and dependency issues surface as confusing runtime errors deep into the migration, at a point where it is expensive to trace them back to a missing package.xml update or an unmet ROS 2 release for a third-party library. Do the audit and build-system conversion first, always.

Step 2: Configure the Operating System for Real-Time Behavior

ROS 2’s real-time story depends on the underlying operating system as much as it depends on the middleware. Most production deployments that need deterministic timing run a real-time kernel patch on Linux.

2.1 Apply a PREEMPT_RT Kernel

The standard approach is to build a Linux kernel with the Fully Preemptible Kernel (PREEMPT_RT) configuration enabled. This changes how the kernel handles interrupt handling and scheduling so that high-priority tasks preempt lower-priority ones with bounded latency, instead of waiting on non-preemptible kernel sections. The ROS real-time working group maintains a kernel builder project specifically for building and configuring RT kernels for ROS 2 testing, which is a far more reliable starting point than hand-patching a stock kernel.

2.2 Isolate CPU Cores for Time-Critical Threads

A real-time kernel alone does not guarantee determinism if time-critical threads compete with general-purpose OS housekeeping for the same cores. Use kernel boot parameters to isolate one or more CPU cores (isolcpus, along with IRQ affinity tuning) and pin the control loop or safety-critical node threads to those isolated cores. Everything else — logging, telemetry upload, the web dashboard — runs on the remaining cores.

2.3 Set Thread Scheduling Policies and Priorities Explicitly

Linux defaults to the CFS (Completely Fair Scheduler) for regular processes. Real-time threads should be moved to SCHED_FIFO or SCHED_RR with an explicit priority, set either via the node’s launch configuration or programmatically at thread creation. Leaving this to OS defaults means the scheduler will treat your control loop the same as a background log-rotation script.

2.4 Budget Memory Up Front

Page faults are a common source of latency spikes in long-running robotic processes. Lock process memory with mlockall and pre-allocate/pre-fault the memory a node will need at startup, rather than letting the allocator page it in during a critical control cycle.

Figure: Layered real-time configuration stack

A production ROS 2 robot’s determinism is the product of four stacked layers working together: a PREEMPT_RT kernel at the base, CPU isolation and IRQ affinity above it, explicit SCHED_FIFO thread priorities above that, and memory locking at the application layer. Weakening any single layer reintroduces jitter regardless of how well the other three are tuned.

Step 3: Tune DDS and QoS for Reliability and Latency

DDS is the communication backbone underneath ROS 2’s rclcpp and rclpy APIs, abstracted through the ROS Middleware (rmw) interface so that different DDS vendors can be swapped without changing application code. Tuning it is a systematic exercise across several dimensions: OS kernel parameters, network topology, message size distribution, and the QoS policy assigned to each topic.

3.1 Choose Reliability Policy Per Topic, Not Globally

ROS 2 defaults to reliable delivery with volatile durability and keep-last history for most topics. That default is a reasonable starting point, but it is wrong for at least two common cases. High-frequency sensor streams, like raw LiDAR point clouds or camera frames, often benefit from best-effort delivery: if a frame is dropped under network congestion, retransmitting a stale frame is worse than skipping it and moving to the next one. Safety-relevant commands and state transitions, on the other hand, need reliable delivery with a bounded depth so that consumers cannot silently miss a critical message.

3.2 Use Durability Deliberately for Late-Joining Subscribers

Transient local durability makes the publisher responsible for persisting the last sample so that a late-joining subscriber — for example, a monitoring dashboard that connects after the robot has already started — receives the current state immediately instead of waiting for the next publish cycle. Volatile durability, the default, does not do this. Static configuration topics and current-state topics are good candidates for transient local; noisy telemetry streams are not.

3.3 Set History Depth Based on Consumer Processing Rate

The history and depth policies together replace the queue-size parameter from ROS 1. Keep-last with a shallow depth (as small as 1 for control commands) prevents a slow consumer from processing a backlog of stale messages. Keep-all is rarely appropriate outside of logging and bag-recording nodes, since it defers all resource limiting to the underlying middleware’s own buffers, which can silently exhaust memory under sustained load.

3.4 Use Deadline, Liveliness, and Lifespan for Fault Detection

Beyond reliability and durability, DDS exposes deadline (flag if a publisher misses an expected publishing interval), liveliness (detect if a publisher has stopped announcing itself), and lifespan (expire samples that have aged past their useful window) policies. These are underused in most ROS 2 deployments, but they are exactly the mechanism that should back a fleet health-monitoring system: a missed deadline on a heartbeat topic is a far more precise fault signal than inferring failure from an absence of expected behavior elsewhere in the stack.

3.5 Tune the Transport Layer for Your Network

Most DDS implementations default to UDP multicast for discovery, which works well on a flat local network but degrades badly across Wi-Fi mesh networks, VLANs, or cloud-connected fleets. Production deployments typically need to either configure discovery servers (to avoid multicast storms on large robot counts) or move to unicast discovery with an explicit peer list. Fragmentation settings also matter: large messages like point clouds get fragmented at the transport layer, and a poorly sized fragment can multiply packet loss sensitivity on lossy Wi-Fi links.

Topic typeRecommended reliabilityRecommended durabilityRecommended depth
Raw sensor stream (LiDAR, camera)Best effortVolatileShallow (1 to 5)
Velocity or motor commandsReliableVolatileVery shallow (1)
Static configuration or mapReliableTransient local1
Diagnostics and heartbeatReliableVolatileSmall, with deadline QoS
Logging and bag recordingBest effort or reliable per needVolatileDeeper, bounded

Step 4: Validate Before Trusting It in the Field

Configuration without measurement is guesswork. Before a ROS 2 stack goes into production, validate the real-time and DDS tuning under conditions that resemble the actual deployment, not the lab bench.

4.1 Measure Latency Under Load, Not at Idle

A control loop that meets its timing budget with the robot sitting still and no other traffic on the network tells you very little about behavior when perception, planning, and logging are all running simultaneously. Load-test with representative CPU and network contention and record worst-case latency, not just average latency.

4.2 Test Network Degradation Explicitly

Simulate packet loss, latency spikes, and temporary network partition (a robot briefly losing Wi-Fi behind a metal rack, for example) and confirm the system degrades gracefully rather than silently. QoS deadline and liveliness settings should surface these events as detectable faults, not as mysterious behavior changes.

4.3 Soak-Test for Memory and Resource Leaks

Run the full stack for many hours, ideally days, under representative load and watch memory, file descriptor counts, and DDS discovery traffic over time. Some DDS implementations exhibit slow discovery-traffic growth on networks with many nodes joining and leaving, which is invisible in a short test but becomes a real problem after a week of continuous fleet operation.

What worked

One deployment team run into DDS discovery storms once a fleet grew past roughly 30 robots on a shared subnet. Rather than reflexively increasing hardware, they moved from default multicast discovery to a discovery server configuration with a small number of designated discovery nodes, and split the fleet into structural DDS domains by function (navigation, perception, fleet management). Discovery traffic dropped sharply and latency on safety-relevant topics became far more consistent, because those topics were no longer fighting for bandwidth against unrelated discovery chatter from unrelated robots.

Common Production Pitfalls

  • Treating QoS defaults as safe. The ROS 2 defaults are reasonable for development but are rarely optimal for every topic in a production system, especially for high-rate sensor streams and safety-relevant commands.
  • Skipping the real-time kernel because “it works fine without it.” Timing that looks fine in a quiet lab environment frequently falls apart once the robot is running with full sensor load, logging, and network traffic simultaneously.
  • Ignoring discovery traffic at fleet scale. Default multicast discovery that works for five robots on a bench does not necessarily scale to fifty robots on a warehouse floor.
  • Migrating logic before build tooling. Runtime bugs introduced by an incomplete migration are far more time-consuming to trace than a build failure caught immediately.
  • No soak testing. Memory growth and discovery-traffic accumulation are slow-burn problems that a short test cannot reveal.
  • rmw layerThe abstraction that lets ROS 2 swap DDS vendors without changing application code — worth knowing because different vendors have meaningfully different default tuning behavior.
  • Domain ID isolationAssigning distinct DDS domain IDs prevents unrelated robots or test rigs on the same network from accidentally discovering and cross-talking with each other.
  • Lifecycle nodesManaged nodes with explicit configured, active, and inactive states make startup ordering and fault recovery predictable instead of ad hoc.
  • Fragmentation sizeLarge messages get split at the transport layer; an oversized fragment on a lossy link multiplies the chance that losing one packet forces a full message retransmit.
  • Colcon build isolationBuilding with colcon’s isolated workspace approach surfaces missing dependencies early, before they become confusing runtime failures.

Key Takeaways

Key Takeaways

  • ROS 2 production readiness depends on deliberate configuration of DDS, QoS, and the real-time kernel — none of it comes correctly set by default.
  • Treat a ROS 1 to ROS 2 migration as a dependency and build-system audit first, and a logic port second.
  • The ros1_bridge package enables incremental migration, letting teams avoid a risky full-stack rewrite.
  • Real-time behavior requires a PREEMPT_RT kernel, CPU isolation, explicit thread scheduling, and memory locking working together as a stack.
  • QoS policy should be set per topic based on its role, not applied globally — sensor streams, commands, and static configuration all need different settings.
  • Deadline and liveliness QoS policies are an underused but precise mechanism for fleet-level fault detection.
  • Load testing, network degradation testing, and multi-day soak testing are all necessary before trusting a ROS 2 stack in the field.

FAQs

Is ROS 2 actually real-time, or just real-time capable?

ROS 2 itself does not guarantee real-time behavior out of the box; it provides the building blocks — a real-time-friendly middleware abstraction and QoS controls — that let engineers achieve deterministic timing when combined with a PREEMPT_RT kernel, CPU isolation, and correct thread scheduling. Without that additional configuration, timing remains best-effort.

Do I need to migrate everything from ROS 1 at once?

No. The ros1_bridge package allows ROS 1 and ROS 2 nodes to exchange messages during a transition, so most teams migrate incrementally, starting with leaf nodes that have few dependents before moving to coordination and planning layers that other nodes depend on.

What is the single most impactful DDS setting to tune first?

Per-topic reliability policy usually has the largest practical impact, since it directly determines whether high-rate sensor streams tolerate drops gracefully and whether safety-relevant commands are guaranteed delivery. Getting this wrong either wastes bandwidth on unnecessary retransmits or silently drops critical messages.

Why does discovery traffic become a problem at fleet scale?

Default DDS discovery relies on multicast announcements between all participants, which scales acceptably for a handful of nodes but generates a growing volume of background traffic as robot count increases, competing with real application traffic. Discovery servers or domain segmentation resolve this by limiting who needs to discover whom.

Do I need a real-time kernel for every ROS 2 robot?

Not necessarily. A robot performing only high-level task coordination with generous timing margins may run fine on a standard kernel. Anything with a tight control loop, direct actuator control, or safety-relevant timing requirements benefits substantially from a PREEMPT_RT kernel and the associated CPU and scheduling configuration.

How is QoS in ROS 2 different from queue size in ROS 1?

ROS 1’s queue size was a single number controlling buffering per subscriber. ROS 2 splits this into separate history and depth policies, plus independent reliability, durability, deadline, liveliness, and lifespan policies, giving far more granular control at the cost of needing to actually configure it deliberately per topic.

What is the biggest real-world source of intermittent latency spikes?

Page faults from unlocked, non-preallocated memory are one of the most common causes, along with real-time threads sharing CPU cores with non-real-time housekeeping processes. Both are addressed by memory locking and CPU isolation respectively.

Can I mix different DDS vendor implementations across a single fleet?

Technically the rmw abstraction allows different DDS vendors to interoperate at the API level, but in practice mixing vendors across a fleet introduces interoperability risk around vendor-specific QoS extensions and discovery behavior. Most production deployments standardize on a single DDS vendor across the fleet to avoid this class of bug entirely.

How long should a soak test run before trusting a ROS 2 stack in production?

There is no universal number, but multi-day continuous operation under representative load is a common minimum, since memory growth and discovery-traffic accumulation are slow-burn effects that short tests will not reveal. Fleets with robots that run for weeks between reboots should soak-test for a comparable duration where feasible.

Glossary

DDS (Data Distribution Service)
A decentralized publish-subscribe communication standard that serves as ROS 2’s underlying middleware, removing the single-master architecture used in ROS 1.
QoS (Quality of Service)
A set of configurable policies in ROS 2, including reliability, durability, history, depth, deadline, liveliness, and lifespan, that control how messages are delivered between publishers and subscribers.
PREEMPT_RT
A Linux kernel patch set that makes the kernel fully preemptible, enabling bounded-latency scheduling for real-time threads.
rmw (ROS Middleware)
The abstraction layer in ROS 2 that allows different DDS vendor implementations to be used interchangeably without changing application code.
Lifecycle node
A ROS 2 node that exposes a managed state machine (unconfigured, inactive, active, finalized), enabling predictable startup ordering and fault recovery.
Discovery server
A DDS configuration that centralizes participant discovery to reduce multicast traffic, used in place of default peer-to-peer multicast discovery on large fleets.

For teams evaluating the broader hardware side of an autonomy stack, our companion piece on choosing sensors for autonomy covers the LiDAR, radar, and camera trade-offs that feed directly into the sensor topics discussed above. Fleet operators managing update rollouts across many ROS 2 robots should also see our guide on over-the-air updates for robot fleets without breaking safety. For the software layer above communication and timing, see our deep dive on the robot perception stack. Teams standardizing on formal safety processes alongside this middleware work may also want our explainer on functional safety for robots, and organizations comparing standards across the whole robotics landscape can consult our overview of the robotics standards landscape.

References

  • ROS 2 Documentation, “Quality of Service settings,” docs.ros.org
  • ROS 2 Design, “ROS 2 Quality of Service policies,” design.ros2.org
  • ROS 2 Design, “ROS QoS – Deadline, Liveliness, and Lifespan,” design.ros2.org
  • ROS 2 Documentation, “Migrating from ROS 1 to ROS 2,” docs.ros.org
  • ros-realtime, “linux-real-time-kernel-builder,” GitHub
  • Embedded Computing Design, “Tune ROS 2 for Commercial Use”
  • Oreate AI Blog, “ROS 2 Guide: Detailed Explanation of DDS Communication Tuning Techniques”
  • Ekumen, “Migrating from ROS 1 to ROS 2: what you need to know”
  • Robotics and Automation News, “ROS 2 and the shift to production robotics,” April 2026
    Camila Duarte
    Camila earned a B.S. in Computer Engineering from Universidade de São Paulo and a postgraduate certificate in IoT Systems from the University of Twente. Her early career took her across farms deploying resilient sensor networks and pushing OTA updates over patchy connections. Those field lessons—battery life, antenna placement, graceful failure—show up in her writing. She focuses on IoT reliability, edge analytics, and sustainability, showing how tiny firmware changes can save energy at scale. Camila co-organizes meetups for women in embedded systems, guest-hosts climate-tech podcasts, and publishes teardown notes of devices that claim to be “low power.” Away from work, she surfs small breaks, does street photography in early light, and hosts feijoada dinners where conversations inevitably drift to UART pins.

      Leave a Reply

      Your email address will not be published. Required fields are marked *