Skip to content

[FIX][TIRx][CUDA] Support low-precision, tensor-map and cluster kernels in CUDA-host bundles - #20489

Merged
MasterJH5574 merged 1 commit into
apache:v0.27.0from
spectrometerHBH:cuda-host-cxx-fixes
Sep 29, 2026
Merged

MasterJH5574 merged 1 commit into
apache:v0.27.0from
spectrometerHBH:cuda-host-cxx-fixes

Conversation

@spectrometerHBH

@spectrometerHBH spectrometerHBH commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

A CUDA-host bundle (tvm.backend.cuda.export_cuda_host, #20395) compiles device and host code as one NVCC C++ translation unit. Building the native kernels of mlc-ai/TIRx-kernels this way exposed several patterns that the bundle could not compile or launch:

  • Low-precision types in the host pass. The CUDA device header included cuda_fp16.h, cuda_bf16.h, cuda_fp8.h, cuda_fp6.h and cuda_fp4.h (and defined the fp8_e4_t-style aliases) only under defined(__CUDA_ARCH__). NVCC's host pass then fails on kernel signatures that use half or nv_bfloat16. The guards now also admit the host pass (!defined(__CUDA_ARCH__) || __CUDA_ARCH__ >= N); NVRTC and device-only NVCC compilation always define __CUDA_ARCH__ and are unchanged.
  • Tensor-map parameters. The host wrapper bound a T.TensorMap() parameter as CUtensorMap* x = ((void*)...), which C accepts and C++ rejects. It now casts explicitly.
  • Argument types at the launch. Host C code spells some device types differently (bfloat16 is uint16_t* on the host, nv_bfloat16* in the kernel), so the direct <<<>>> launch did not type-check. Launches now go through a small helper that converts each argument to the kernel's own parameter type and calls cudaLaunchKernelEx.
  • Launch attributes. clusterCtaIdx.*, preferredClusterCtaIdx.*, programmatic dependent launch and cooperative launch were rejected. They now become the corresponding cudaLaunchAttributes, following CUDAWrappedFunc in the CUDA runtime module (including cudaFuncAttributeNonPortableClusterSizeAllowed for cluster launches and omitting a unit preferred cluster). Required block dimensions remain unsupported and are still diagnosed.

Testing

On B200 (sm_100a), CUDA 13.2:

  • tests/python/codegen/test_target_codegen_cuda.py -k cuda_host: 8 passed (4 tests, each under NVCC and NVRTC device compilation). New tests:
    • bfloat16 buffers with a two-CTA cluster launch: runs and checks both the values and each CTA's cluster rank ([0, 1, 0, 1]);
    • a programmatic dependent launch flag: runs and checks the launch attribute;
    • a tensor-map parameter: compiles the bundle with NVCC.
  • tests/python/codegen/test_target_codegen_cuda.py, tests/python/tirx/codegen/test_codegen_cuda.py, tests/python/tirx-transform/test_tir_transform_split_host_device.py, tests/python/tirx-transform/test_tir_transform_lower_tvm_builtin.py: 614 passed, 8 skipped.
  • End to end with mlc-ai/TIRx-kernels (with fix(kernel): encode tensor maps with the typed op for cuda_host builds mlc-ai/TIRx-kernels#3, which switches its hand-written packed tensor-map encodes to tensormap_encode_tiled): every CUDA compile was built with host="cuda_host", bundled with export_cuda_host, compiled by tvm_ffi.cpp.build_inline and loaded with tvm_ffi.load_module. All 220 correctness configs of fp16_bf16_gemm, nvfp4_gemm, rmsnorm, alphamoe_fp8_blockscale_qwen3next, kda_decode_multishape, msa_prefill_multishape, msa_decode_multishape, vsa_multishape and mla_dsv4_multishape pass. Before these two changes, none of these kernels built as a bundle.

The same change for main is #20490.

…ls in CUDA-host bundles

A CUDA-host bundle compiles device and host code as one NVCC C++
translation unit. Several kernel patterns failed there:

- The low-precision headers and type aliases emitted for CUDA device code
  were guarded by __CUDA_ARCH__, so NVCC's host pass could not parse kernel
  signatures that use half, nv_bfloat16 or the fp8/fp6/fp4 types. Expose
  them to the host pass as well; NVRTC and device-only NVCC compilation are
  unchanged.
- Tensor-map parameters were bound through C's implicit void* conversion,
  which C++ rejects. Cast to CUtensorMap* explicitly.
- Host C types spell some device types differently (bfloat16 is uint16_t),
  so direct <<<>>> launches did not type-check. Launch through a helper
  that converts each argument to the kernel's parameter type and calls
  cudaLaunchKernelEx.
- Cluster, preferred-cluster, programmatic dependent and cooperative launch
  parameters were rejected. Emit the matching launch attributes, following
  the CUDA runtime module. Required block dimensions remain unsupported.
@MasterJH5574
MasterJH5574 merged commit 7b5acb3 into apache:v0.27.0 Sep 29, 2026
0 of 5 checks passed
tqchen pushed a commit that referenced this pull request Sep 29, 2026
…ls in CUDA-host bundles (#20490)

A CUDA-host bundle (`tvm.backend.cuda.export_cuda_host`, #20395)
compiles device and host code as one NVCC C++ translation unit. Building
the native kernels of mlc-ai/TIRx-kernels this way exposed several
patterns that the bundle could not compile or launch:

- **Low-precision types in the host pass.** The CUDA device header
included `cuda_fp16.h`, `cuda_bf16.h`, `cuda_fp8.h`, `cuda_fp6.h` and
`cuda_fp4.h` (and defined the `fp8_e4_t`-style aliases) only under
`defined(__CUDA_ARCH__)`. NVCC's host pass then fails on kernel
signatures that use `half` or `nv_bfloat16`. The guards now also admit
the host pass (`!defined(__CUDA_ARCH__) || __CUDA_ARCH__ >= N`); NVRTC
and device-only NVCC compilation always define `__CUDA_ARCH__` and are
unchanged.
- **Tensor-map parameters.** The host wrapper bound a `T.TensorMap()`
parameter as `CUtensorMap* x = ((void*)...)`, which C accepts and C++
rejects. It now casts explicitly.
- **Argument types at the launch.** Host C code spells some device types
differently (bfloat16 is `uint16_t*` on the host, `nv_bfloat16*` in the
kernel), so the direct `<<<>>>` launch did not type-check. Launches now
go through a small helper that converts each argument to the kernel's
own parameter type and calls `cudaLaunchKernelEx`.
- **Launch attributes.** `clusterCtaIdx.*`, `preferredClusterCtaIdx.*`,
programmatic dependent launch and cooperative launch were rejected. They
now become the corresponding `cudaLaunchAttribute`s, following
`CUDAWrappedFunc` in the CUDA runtime module (including
`cudaFuncAttributeNonPortableClusterSizeAllowed` for cluster launches
and omitting a unit preferred cluster). Required block dimensions remain
unsupported and are still diagnosed.

### Testing

On B200 (sm_100a), CUDA 13.2, this branch:

- `tests/python/codegen/test_target_codegen_cuda.py -k cuda_host`: 8
passed (4 tests, each under NVCC and NVRTC device compilation). New
tests:
- bfloat16 buffers with a two-CTA cluster launch: runs and checks both
the values and each CTA's cluster rank (`[0, 1, 0, 1]`);
- a programmatic dependent launch flag: runs and checks the launch
attribute;
  - a tensor-map parameter: compiles the bundle with NVCC.
- `tests/python/codegen/test_target_codegen_cuda.py`,
`tests/python/tirx/codegen/test_codegen_cuda.py`,
`tests/python/tirx-transform/test_tir_transform_split_host_device.py`,
`tests/python/tirx-transform/test_tir_transform_lower_tvm_builtin.py`:
487 passed, 8 skipped. The remaining 120 failures are all
parametrizations of `test_ptx_cp_async`, which fails while parsing its
TVMScript body (`prim._OpEQ` receives a string), before any code
generation; this change does not touch that path.

The same change is proposed for the v0.27.0 release branch in #20489.
There it was also verified end to end with mlc-ai/TIRx-kernels (whose
kernels currently target the v0.27.0 script APIs): with
mlc-ai/TIRx-kernels#3, all 220 correctness configs of its nine
single-GPU native kernels pass when every CUDA compile is built with
`host="cuda_host"`, bundled with `export_cuda_host`, compiled by
`tvm_ffi.cpp.build_inline` and loaded with `tvm_ffi.load_module`.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants