Skip to content

perf: cut ROS2 ingest CPU ~2x, decline WS deflate; Lyrical CI; dependency bumps - #13

Merged
facontidavide merged 9 commits into
mainfrom
perf/cpu-fixes
Sep 20, 2026
Merged

facontidavide merged 9 commits into
mainfrom
perf/cpu-fixes

Conversation

@facontidavide

@facontidavide facontidavide commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes driven by real profiling (perf + pidstat against ros2 bag play of real bags); analysis and numbers are in the companion docs PR (#14, docs/CPU_OPTIMIZATION.md §2.0 / §2.9).

  • ROS2 ingest rework. >50% of process CPU was rcl_wait rebuilding the wait set once per received message. Subscriptions now drain their reader (GenericSubscription subclass), the ingest executor polls every ingest_poll_interval_ms (default 5, 0 = old blocking spin, blocks while there are no subscriptions so idle is unchanged), and publish/request/timeout timers run on their own executor thread so zstd never blocks ingest. min_qos_depth default 1 → 10. Receive timestamps now come from the RMW received_timestamp.
  • Decline WebSocket permessage-deflate. IXWebSocket re-deflated our zstd frames for any client offering the extension (browsers, Python websockets; the PlotJuggler plugin does not offer it).
  • Ingest/serializer copies no longer zero the buffer first.
  • MessageBuffer::set_latched() seeds from an already-buffered sample (race exposed by threaded ingest; failing test first).
  • Lyrical: target_link_libraries instead of ament_target_dependencies (removed in Lyrical), new ROS: Lyrical CI workflow, Lyrical .deb in the release matrix. No pixi/conda/AppImage for Lyrical: there is no robostack-lyrical channel.
  • Dependencies: IXWebSocket 11.4.6 → 12.0.1 (FetchContent + .deb build; fixes only, several server-side), Fast DDS 3.4.0 → 3.4.3, CLI11 2.6.0 → 2.6.2. Fixed a pre-existing link failure of the standalone backends when IXWebSocket is fetched with TLS. pixi.lock intentionally untouched: moving the Humble env to IXWebSocket 12 re-solves the whole RoboStack stack — separate chore.

Measurements

i7-13700H, bridge pinned to P-cores, performance governor, 1 client, losslessness verified by per-topic message counts decoded client-side.

Workload main this PR
83 topics, ~700 msg/s small messages 7.2–7.6% of a core 3.4–3.7%
4 lidars 52 MiB/s + 400 Hz imu + 480 Hz tf 17.9–20.2% 18.8–19.4% (zstd-bound, unchanged)
bag playing, no clients 0.45% 0.45%
client offering permessage-deflate (unpinned rig) 19.2% 12.9%

Dead end worth knowing: polling without draining silently loses data (executor takes one message per subscription per wait cycle; 481 Hz /tf fell to 154 Hz).

Test plan

  • 271 unit tests pass on Humble (new: deflate handshake, drain-in-one-pass, latched seeding), also against IXWebSocket 12.0.1 via FetchContent
  • Builds on Jazzy; subscription-manager suites (25 tests) pass there
  • Builds on system Lyrical; live end-to-end run lossless (29798/29798 msgs). ROS unit suites not runnable on my Lyrical install (tests force rmw_cyclonedds_cpp, not installed)
  • Live ThreadSanitizer run of the two-thread layout under lidar load with clients joining/leaving: no races in pj_bridge code (run before the review-fix commit; not repeated after)
  • Codex review; findings addressed in c57a313
  • FastDDS backend builds and starts with fast-dds 3.4.3 / CLI11 2.6.2 / IXWebSocket 12.0.1
  • Lyrical .deb job runs only on tags / manual dispatch — not exercised by this PR
  • CI on all distros, incl. the new Lyrical job

🤖 Generated with Claude Code

facontidavide and others added 8 commits September 20, 2026 12:43
Binary frames are already zstd-compressed. When a client offered the
extension, IXWebSocket deflated every frame again with zlib on the publish
thread: 7-11% of process CPU in a perf profile, 19.2% -> 12.9% of a core
for one client on an 83-topic workload.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Profiling showed >50% of process CPU in rcl_wait: rmw_fastrtps re-attaches
every subscription to the wait set on each wait cycle, and blocking spin()
does one cycle per received message.

- Subscription callbacks now drain their reader, so one cycle harvests a
  whole burst (the executor hands over only one message per cycle).
- The ingest executor polls every ingest_poll_interval_ms (default 5,
  0 = old blocking spin).
- min_qos_depth default 1 -> 10 so a poll interval cannot overflow a
  shallow reader.
- Publish/request/timeout timers moved to a dedicated callback group and
  executor thread: zstd no longer blocks ingest.

83 topics, ~700 msg/s, 1 client: 7.5% -> 3.4% of a core, identical
message counts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…st poll

- heavy_frame_zstd_level (ROS2 param / --heavy-frame-zstd-level): zstd is
  ~60% of CPU on point-cloud workloads. Default stays 1 (no behaviour
  change); measured on 4 lidars, 52 MiB/s in: 1 = 19% of a core, 31 MB/s
  out; -5 = 15%, 39 MB/s; -100 = 10%, 51 MB/s. Wire-compatible.
- Ingest and serializer copies no longer zero the buffer first.
- ROS2 ingest loop blocks instead of polling while there are no
  subscriptions, so idle CPU matches the old blocking spin (0.45%).
- Validation, docs and changelog for the new options.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- Drain moved into a GenericSubscription subclass (handle_serialized_message)
  instead of a callback holding a weak_ptr to its own subscription: the
  pointer was assigned after the subscription became executable, racing with
  the ingest thread.
- Receive time now comes from the RMW received_timestamp when available;
  with batched delivery "now" would cluster timestamps and distort per-client
  rate limiting.
- MessageBuffer::set_latched() seeds the retained sample from the buffer:
  with ingest on its own thread the sole latched sample can arrive before
  the subscriber's bookkeeping, and later clients would get no replay.
- Timer thread is joined on every exit path and reports exceptions instead
  of terminating; take errors during drain are caught.

Builds and passes subscription tests on Humble and Jazzy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
ament_target_dependencies was removed in newer distros, so the package did
not configure on Lyrical. Modern targets work on Humble, Jazzy and Lyrical.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
No robostack-lyrical channel exists, so there is no pixi/conda/AppImage
variant for Lyrical; colcon CI and the Debian release are covered.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
All frames keep zstd level 1 as before; the knob was a CPU-vs-bandwidth
tradeoff not worth the configuration surface. FastDDS/RTI entry points are
back to identical with main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…LS link

IXWebSocket 12.x is fixes only (server header parsing, fd double close,
DNS crash, move instead of copy in send buffering). pixi.lock is left
alone: moving the Humble env to 12.x re-solves the whole RoboStack stack.

The fetched static IXWebSocket does not export its OpenSSL dependency, so
the standalone FastDDS/RTI link failed with TLS on (also on 11.4.6).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@facontidavide facontidavide changed the title perf: cut ROS2 ingest CPU ~2x, decline WS deflate, heavy-frame zstd knob perf: cut ROS2 ingest CPU ~2x, decline WS deflate; Lyrical CI; dependency bumps Sep 20, 2026
…exit

On Lyrical, rcl_logging_spdlog registers a periodic flush thread on the
global spdlog registry and rcl unloads the library on shutdown. After the
suites' repeated init/shutdown the thread resumed in unmapped code at exit
(all tests passed, process returned -11). The bridge binary is unaffected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@facontidavide
facontidavide merged commit ba4ea90 into main Sep 20, 2026
7 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant