Skip to content

Fixed-size reductions - #1061

Merged
cliffburdick merged 2 commits into
mainfrom
cub_segmented_same
Sep 30, 2025
Merged

cliffburdick merged 2 commits into
mainfrom
cub_segmented_same

Conversation

@cliffburdick

Copy link
Copy Markdown
Collaborator

Add optimized fixed-size segmented reduce support for CUB 3.2+

Implement conditional support for CUB's new fixed-size segmented reduction API
when using CUB version 3.2 or higher. This optimization applies to both generic
reduce and sum operations when:

  • Input is a tensor view
  • Both input and output tensors are contiguous memory
  • All segments have uniform size

The new implementation uses CUB's simplified API that takes a segment size
parameter instead of begin/end iterators, providing better performance for
uniform segment reductions.

Changes:

  • Add version checks for CUB_MAJOR_VERSION >= 3 && CUB_MINOR_VERSION >= 2
  • Implement fast path using fixed segment size for contiguous tensors
  • Fall back to iterator-based API for non-contiguous or non-uniform cases
  • Apply optimization to both DeviceSegmentedReduce::Reduce and ::Sum

This maintains backward compatibility while leveraging newer CUB features
when available.

@copy-pr-bot

copy-pr-bot Bot commented Sep 29, 2025

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@cliffburdick

Copy link
Copy Markdown
Collaborator Author

/build

@cliffburdick
cliffburdick merged commit c01118e into main Sep 30, 2025
1 check passed
@cliffburdick
cliffburdick deleted the cub_segmented_same branch September 30, 2025 18:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant