Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 42 additions & 0 deletions docs/release-highlights.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,48 @@ How it was measured:

The results describe this workload, not general broker capacity: run the scenario on your own hardware to compare with your application. Consumers still need enough processing capacity to keep up with producers.

### Seek by timestamp in offloaded topics with fewer object store requests {#seek-by-timestamp-in-offloaded-topics}

Topics that keep a long history, such as logs retained for a year, hold most of it in [tiered storage](tiered-storage-overview.md). When a consumer seeks to a timestamp in that history, the broker runs a binary search that reads one entry per step, and the index of an offloaded ledger records only the first entry of each data block. Reaching any other entry meant scanning the block from its start, one ranged read at a time. Pulsar 5.0 reduces that scanning, so that a search sends fewer requests to the object store and downloads less data.

A benchmark counted the ranged reads that searches by timestamp send to the object store, before and after three changes to how the jclouds offloader and the search read offloaded ledgers:

- A topic holds about **130,000 entries of 2 to 6 KiB** in **two offloaded ledgers of about 250 MiB each**, with 64 MiB data blocks and a 1 MiB read buffer.
- A **cold search** runs right after the topic is loaded, with no entry offset cached.
- A **warm search** follows each cold search and looks for a position 200 entries earlier, as when a consumer seeks again nearby.

The results:

<div className="release-highlights-benchmark">

| Ranged reads per search | Before | 5.0.0 | Change |
| --- | ---: | ---: | ---: |
| **Cold search, mean of 16 searches** | **428** | **44** | **−90%** |
| Warm search, mean of 16 searches | 223 | 9 | −96% |
| Cold search, maximum | 738 | 70 | −91% |

</div>

Each ranged read requests 1 MiB, so the same figures are the MiB downloaded per search. **Before** is the code without any of the three changes below, as in Pulsar 4.0.13 and 4.2.4. The first change is also in the 4.0 and 4.2 maintenance branches for their next releases.

What each change contributes:

- **Entry offsets found while scanning are reused:** a scan caches the offset of every entry it passes, and a later read starts from a nearby cached offset instead of the start of the block. Cold searches go from 428 to 156 ranged reads, and warm searches from 223 to 14.
- **The search reads the first entry of data blocks:** when the first entry of a data block is close to the midpoint of the search range, the search reads that entry instead, with a single ranged read. Cold searches go from 156 to 61, and warm searches from 14 to 9.
- **Offsets are remembered for each open ledger:** scans keep offsets about 1 MiB apart for as long as the offloaded ledger stays open, and later reads in the same data block resume from the nearest one. Cold searches go from 61 to 44, and warm searches stay at 9.

Ranged reads per search before the changes and after each change, cold and warm searches each in a panel of their own with the same axis:

![Ranged reads sent to the object store per search by timestamp on offloaded ledgers, before the changes and after each of the three changes, for cold and warm searches in separate panels with the same axis](/assets/release-highlights-5.0/seek-by-timestamp-offloaded-ranged-reads.svg)

How it was measured:

- [`OffloadedLedgerFindPositionTest`](https://github.com/apache/pulsar/blob/master/tiered-storage/jcloud/src/test/java/org/apache/bookkeeper/mledger/offload/jcloud/impl/OffloadedLedgerFindPositionTest.java) ran before the first change and after each of the three changes, with its data set to 64 MiB data blocks and two offloaded ledgers of 64,997 entries. As committed, the test uses 16 MiB data blocks and one ledger to stay small.
- The object store was in memory, so the benchmark counts requests and bytes, not latency. The broker entry cache was disabled so that every entry read reached the offloaded ledger.
- The 16 searched positions are spread over both ledgers and over the positions within their data blocks. Each search covers both ledgers; a seek first narrows the search to the ledgers that can contain the timestamp.

The results describe this data layout, not seek latency: each ranged read is a request to your object store, and its latency and cost depend on that store. [Read performance for object storage](tiered-storage-overview.md#read-performance-for-object-storage) describes the settings involved, and the [message position search metrics](reference-metrics-opentelemetry.md#message-position-search-metrics) show how long searches take and how much they read on a running broker.

### Less work per message {#less-work-per-message}

The efficiency improvements reach across the messaging path:
Expand Down
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
42 changes: 42 additions & 0 deletions versioned_docs/version-5.0.x/release-highlights.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,48 @@ How it was measured:

The results describe this workload, not general broker capacity: run the scenario on your own hardware to compare with your application. Consumers still need enough processing capacity to keep up with producers.

### Seek by timestamp in offloaded topics with fewer object store requests {#seek-by-timestamp-in-offloaded-topics}

Topics that keep a long history, such as logs retained for a year, hold most of it in [tiered storage](tiered-storage-overview.md). When a consumer seeks to a timestamp in that history, the broker runs a binary search that reads one entry per step, and the index of an offloaded ledger records only the first entry of each data block. Reaching any other entry meant scanning the block from its start, one ranged read at a time. Pulsar 5.0 reduces that scanning, so that a search sends fewer requests to the object store and downloads less data.

A benchmark counted the ranged reads that searches by timestamp send to the object store, before and after three changes to how the jclouds offloader and the search read offloaded ledgers:

- A topic holds about **130,000 entries of 2 to 6 KiB** in **two offloaded ledgers of about 250 MiB each**, with 64 MiB data blocks and a 1 MiB read buffer.
- A **cold search** runs right after the topic is loaded, with no entry offset cached.
- A **warm search** follows each cold search and looks for a position 200 entries earlier, as when a consumer seeks again nearby.

The results:

<div className="release-highlights-benchmark">

| Ranged reads per search | Before | 5.0.0 | Change |
| --- | ---: | ---: | ---: |
| **Cold search, mean of 16 searches** | **428** | **44** | **−90%** |
| Warm search, mean of 16 searches | 223 | 9 | −96% |
| Cold search, maximum | 738 | 70 | −91% |

</div>

Each ranged read requests 1 MiB, so the same figures are the MiB downloaded per search. **Before** is the code without any of the three changes below, as in Pulsar 4.0.13 and 4.2.4. The first change is also in the 4.0 and 4.2 maintenance branches for their next releases.

What each change contributes:

- **Entry offsets found while scanning are reused:** a scan caches the offset of every entry it passes, and a later read starts from a nearby cached offset instead of the start of the block. Cold searches go from 428 to 156 ranged reads, and warm searches from 223 to 14.
- **The search reads the first entry of data blocks:** when the first entry of a data block is close to the midpoint of the search range, the search reads that entry instead, with a single ranged read. Cold searches go from 156 to 61, and warm searches from 14 to 9.
- **Offsets are remembered for each open ledger:** scans keep offsets about 1 MiB apart for as long as the offloaded ledger stays open, and later reads in the same data block resume from the nearest one. Cold searches go from 61 to 44, and warm searches stay at 9.

Ranged reads per search before the changes and after each change, cold and warm searches each in a panel of their own with the same axis:

![Ranged reads sent to the object store per search by timestamp on offloaded ledgers, before the changes and after each of the three changes, for cold and warm searches in separate panels with the same axis](/assets/release-highlights-5.0/seek-by-timestamp-offloaded-ranged-reads.svg)

How it was measured:

- [`OffloadedLedgerFindPositionTest`](https://github.com/apache/pulsar/blob/master/tiered-storage/jcloud/src/test/java/org/apache/bookkeeper/mledger/offload/jcloud/impl/OffloadedLedgerFindPositionTest.java) ran before the first change and after each of the three changes, with its data set to 64 MiB data blocks and two offloaded ledgers of 64,997 entries. As committed, the test uses 16 MiB data blocks and one ledger to stay small.
- The object store was in memory, so the benchmark counts requests and bytes, not latency. The broker entry cache was disabled so that every entry read reached the offloaded ledger.
- The 16 searched positions are spread over both ledgers and over the positions within their data blocks. Each search covers both ledgers; a seek first narrows the search to the ledgers that can contain the timestamp.

The results describe this data layout, not seek latency: each ranged read is a request to your object store, and its latency and cost depend on that store. [Read performance for object storage](tiered-storage-overview.md#read-performance-for-object-storage) describes the settings involved, and the [message position search metrics](reference-metrics-opentelemetry.md#message-position-search-metrics) show how long searches take and how much they read on a running broker.

### Less work per message {#less-work-per-message}

The efficiency improvements reach across the messaging path:
Expand Down